SOURCE-LINKED INTELLIGENCE
A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems
Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (MLLMs) require different forms of assessment, yet their evaluation evidence must remain interpretable for development and oversight. We propose a unified framework that connects output-level, trajectory-level, and cross-modal assessment through eight trustworthiness dimensions: capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency. The framework preserves syst
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T00:31:16.000Z
- arXiv · Artificial Intelligence · 2026-09-17T00:31:16.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.