New✦ Any source in, training data out
Easily make ML-ready datasets
Bring your data, tell DATAtokeniz how to handle it, and its deterministic pipeline compiles it — with exact provenance on every row.
PDF → graded eval setno key needed
Transformer
Input → guided handling → JSONL outputInput
Source material
PDF, DOCX, slides, CSV, JSON, images, audio, video or a link. Processed in the request and never stored. Audio and video are transcribed into timestamped segments; a public Kaggle dataset is downloaded and every row cites its file.
Declare source license— optional
Demo mode · nothing is saved. Sign in to raise the limits and keep a run history.
Transformer
Tell DATAtokeniz how to handle it
Guide
Add a source on the left, then tell me exactly how to turn it into a dataset.
- extract
- segment
- enrich
- quality
- compile
Extracted only — the guide never writes dataset rows.
Output
ML-ready dataset
Your dataset appears here after a run — preview, quality score and one-click download.