AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence

arXiv · AI, language, vision and robotics · article · Aug 26, 2026 · UTC

Large language models can now generate complex, multi-step mathematical proofs, but reliably determining their correctness and localizing early logical errors remains a critical challenge. Existing evaluation approaches largely depend on model-based natural-language judgments, which often overlook local reasoning gaps. While formal theorem provers like Lean offer a path to rigorous verification, using them to evaluate informal text requires solving locality and semantic mismatches: a prover might bypass a local flaw by proving an overly broad target, or validate an auto-formalized statement th

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T09:11:58.312Z. This is not the publication date.