SOURCE-LINKED INTELLIGENCE
AutoData: Agentic Search for Pre-training Data Selection
LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, str
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T06:24:52.000Z
- arXiv · Artificial Intelligence · 2026-09-17T06:24:52.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.