SOURCE-LINKED INTELLIGENCE
Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning
Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answers with a filter-then-clone pipeline: a state-only value trained from segment comparisons scores whole trajectories, Learn-then-Test calibration certifies a selection threshold under a distribution-free $(α, δ)$ bound on the unsafe fraction of the selection, and behavior cloning follows. We are not aware of prior work certifying the comp
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-10T06:45:32.000Z
First collected: 2026-09-20T19:12:12.556Z. This is not the publication date.