SOURCE-LINKED INTELLIGENCE
PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces
Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific reasoning skills, rely on external knowledge, or are costly to extend. We introduce PetriBench, a compact, fully self-contained, and scalable benchmark for evaluating LLM reasoning over dynamic state spaces using Petri nets, a mature formalism for modeling real-world concurrent and distributed systems. PetriBench organizes reasoning into four task families varying by scope and temporal horizon, with Easy, Medium, and Hard levels generated by increasing structural complexity and evaluated agains
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · Artificial Intelligence · 2026-09-17T08:31:18.000Z
- arXiv · AI, language, vision and robotics · 2026-09-17T08:31:18.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.