The dataset is free.
The audit-ready record isn't.
Anyone can turn a PDF into JSONL. DATAtokeniz proves where every row came from — labelling what was extracted versus generated, and recording the origin of every source, so the paperwork exists before anyone asks for it.
EU AI Act Article 50 transparency obligations and the GPAI training-data summary apply from 2 August 2026. Article 10 data governance for high-risk systems follows on 2 December 2027.
Free
The whole pipeline, no card.
$0forever
500 pages / month
Start free- Every output format — eval, RAG, SFT, multimodal, Parquet
- Quality score and full quality report
- Article 10 record — preview (complete findings, not countersigned)
- Provenance on every row
- With sign-in: save to your own Google Drive and keep a verifiable run history
- Try it with no account at all
Pro
Records you can put in an audit file.
$66/mo · $790 billed yearly
5,000 pages / month
- Everything in Free
- Issued Article 10 records — countersigned and committed to a tamper-evident ledger
- 500 runs / month and 100 MiB files — 25× the runs, 2× the file size of Free
- 14-day free trial — card up front, cancel in one click from the customer portal
Team
One evidence trail for the whole team.
$416/mo · $4,990 billed yearly
25,000 pages / month
Talk to us- Everything in Pro, 5 seats included
- Shared datasets and run history
- Priority processing on the durable lane
- Consolidated billing
Enterprise
From $25k / year.
Custom
Unlimited, by agreement
Talk to us- Everything in Team
- Bring your own signing key — records verify without trusting DATAtokeniz
- VPC or on-premise deployment
- DPA, SLA, SSO and custom retention
- Audit export and named support
Why it isn't priced per page
Parsers sell pages, and pages are a race to zero. DATAtokeniz charges for the thing that survives that race: an attested chain from your source file to the row a model trained on. The parsing is table stakes and we give it away.
| Capability | Document parsers | DATAtokeniz |
|---|---|---|
| Source → clean text | Yes | Yes |
| Training-ready JSONL (eval / RAG / SFT) | Roll your own | Included, free |
| Per-row provenance to page or timestamp | — | Included, free |
| PII & secret detection, secrets never exported | — | Included, free |
| Extracted vs. model-generated labelled on every row | — | Included, free |
| Article 10 data-governance record | — | Preview free · issued on Pro |
| Tamper-evident run ledger | — | Included, with sign-in |
| Verify without trusting the vendor | — | Enterprise |
Questions
What exactly is behind the paywall?
The attestation, never the evidence. Free and Pro produce the same Article 10 findings — origin, licences, preparation operations, bias checks and personal-data findings. A free record is stamped preview inside the body its SHA-256 covers, so anyone reading it learns it was not issued; a Pro record is issued and its run is committed to a tamper-evident ledger you can verify. The stamp cannot be edited out — changing it breaks the digest.
Does this make my AI system compliant with the EU AI Act?
No, and any vendor who tells you otherwise is selling you a risk. DATAtokeniz generates the data-governance documentation Article 10 asks providers to keep, derived from what the pipeline actually observed. It is evidence for a compliance file. It is not legal advice and it is not a conformity assessment — you still need those.
Which EU AI Act deadline actually applies to me?
Two different ones, and they are often conflated. Article 50 transparency obligations — including marking AI-generated content — and the GPAI training-data summary under Annex XI apply from 2 August 2026. Article 10 data governance for high-risk systems applies from 2 December 2027. DATAtokeniz helps with both: it labels every record as directly extracted or model-generated, which is the substance of the transparency duty, and it compiles the Article 10 record for the later one. Confirm your own obligations with counsel — which articles bind you depends on your role and risk class, not on us.
How does a page get counted?
Once per source page, not once per output row. A 40-page PDF that yields 300 records counts as 40 pages. Counts are measured from the finished dataset, so any invoice line can be reconstructed from the dataset itself.
Is there a free trial?
Yes — Pro starts with a 14-day free trial, once per customer. A card is required up front so the subscription can continue without interruption, but cancel from the customer portal before the trial ends and you are charged nothing.
What happens when I hit my limit?
The run is refused until the next period, on every plan — DATAtokeniz never bills you for anything you did not explicitly buy, so there is no overage charge and no surprise invoice. Pro's allowances are 25× the runs and 10× the pages of Free, which is where most growing teams land.
Where is my data stored?
Nowhere DATAtokeniz owns. Inputs are processed in the request and never written to DATAtokeniz storage; results go to your own Google Drive. Drive access is drive.file scope only — files you pick or that DATAtokeniz creates.
Can I verify a record without trusting DATAtokeniz?
On Enterprise, yes — you hold the signing key, so records verify against a secret DATAtokeniz never sees. On every tier the SHA-256 digest is recomputable from the record body, which detects any edit to the findings.
Run it before you pay anything
Three runs, any format, no account. You'll see the quality score and the full provenance record before you decide.