AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Agents

Explore collected AI evidence about Agents, with dates and links to original sources.

Showing 20 of 712 matching collected records. Text matches can include mentions by other organizations.

  1. Sep 19, 2026 · UTC · BleepingComputer

    BragJack attacks hijack AI browser agents through malicious extensions

    BragJack, a proof-of-concept attack from Forever Security's Gal Weizman, hijacks the AI assistants in Chrome, Edge, Opera Neon, Perplexity Comet, and Claude in Chrome using one malicious extension. The Prompt Forcing technique earned over $20,000 in bounties and two CVEs. [...]

  2. Sep 18, 2026 · UTC · AWS Artificial Intelligence Blog

    Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime

    Migrate a multi-model healthcare AI agent from self-managed Amazon ECS with AWS Fargate to Amazon Bedrock AgentCore runtime, preserving triple-model orchestration and vector-enhanced knowledge retrieval while reducing infrastructure management. The framework-agnostic pattern applies across healthcare, financial services, and manufacturing.

  3. Sep 18, 2026 · UTC · AWS Artificial Intelligence Blog

    The new AgentCore runtime: Elastic, optimized, and consistently fast starts

    Today we are announcing the new AgentCore runtime, a capability of Amazon Bedrock AgentCore built for the speed, flexibility, and cost efficiency that production agents demand. It reclaims memory as sessions release it and delivers consistent cold starts regardless of image size or concurrency.

  4. Sep 18, 2026 · UTC · AWS Artificial Intelligence Blog

    Deploy Hugging Face models on Amazon SageMaker AI with coding agents

    Deploy production-ready Hugging Face models on Amazon SageMaker AI using six open-source agent skills. Point a coding agent at a model and get back a real-time endpoint with the right serving container, autoscaling, Amazon CloudWatch alarms, and a verified teardown path.

  5. Sep 18, 2026 · UTC · The Hacker News

    Plugin4Shell Lets Repository Owners Swap Pinned Plugin Code Across Four AI Coding Agents

    A flaw in four widely used AI coding agents lets someone who controls a plugin's code repository swap the plugin an agent installs for a malicious one, even when the agent locked that plugin to a specific reviewed version, security firm Air Security said on Thursday. The firm said Anthropic has patched the flaw in Claude Code 2.1.179 and OpenAI in Codex 0.146.0, that GitHub Copilot has no

  6. Sep 17, 2026 · UTC · The Verge AI

    The AI Superintelligence Slowdown

    Remember when tech leaders would tell their employees to “move fast and break things”? It seemed that would be the way of AI too. But after a summer where rogue AI agents became reality, and researchers warned that AI could kill us all, a number of leading US AI companies are publicly suggesting it’s time […]

  7. Sep 17, 2026 · UTC · BleepingComputer

    OpenAI details more cases of AI agents taking unauthorized actions

    OpenAI has presented new examples of what they call "AI model misalignment" from the past six months, including unauthorized file uploads, following self-generated instructions, hiding mistakes, and leveraging exposed API keys. [...]

  8. Sep 17, 2026 · UTC · arXiv · Artificial Intelligence

    Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation

    Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and agents built in this way now operate robots without robot-specific training.Whether this paradigm is also safe, however, has not been asked. We evaluate coding agent under a safety constraint, where each task pairs a manipulation goal with an obstacle the robot must not touch. The agent pursues the goal but collides with the obstacle in most cases, treating task completion as its sole objective while neglecting safety. The agent reasons about the obstacle in

  9. Sep 17, 2026 · UTC · arXiv · AI, language, vision and robotics

    Quantifying Overclaiming Propensity in Frontier LLM Agents

    Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent's final response is often the only account of that work a user sees. We quantify the propensity of frontier agents to \emph{overclaim} task completion, a misrepresentation that can mislead the user. An agent overclaims when its final response contradicts information in its context. This definition requires no inference about intent and is independent of task success. We introduce \emph{OverclaimBench}, an evaluation suite composed of five file-review scenarios, transcript-based coverage measurem

  10. Sep 17, 2026 · UTC · arXiv · Artificial Intelligence

    An Empirical Study of Harness Design for Coding Agents

    Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five conte

  11. Sep 17, 2026 · UTC · arXiv · Artificial Intelligence

    RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

    Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, ski

  12. Sep 17, 2026 · UTC · arXiv · AI, language, vision and robotics

    Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights

    Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. We present the semantic action graph, a lightweight domain schema that represents a sports match as performer, action, recipient, moment, and state nodes connected by role, temporal, and outcome edges. The schema demonstrates three key properties: 1) connected event sequences, 2) a shared, closed vocabulary, and 3) frame-addressable momen

  13. Sep 17, 2026 · UTC · arXiv · Artificial Intelligence

    Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

    Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. We treat the evaluation set as a finite population and seek accurate point and interval estimates of each domain mean. Direct estimators, including prediction-powered inference (PPI), use only a domain's own labels and are imprecise where labels are few. Small area estimation addresses this problem, and we build on it to develop an integrated workf

  14. Sep 17, 2026 · UTC · arXiv · Artificial Intelligence

    RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents

    Effective troubleshooting agents in enterprise customer support depend on retrieving actionable guidance from similar historical cases, yet existing retrieval-augmented generation (RAG) systems treat support cases as static documents and overlook their multi-stage, stateful nature. We introduce RAFT (Retrieval-Augmented Framework for Troubleshooting Agents), a stateful RAG framework that abstracts each closed historical case into a directed chain of timeline entries and retrieves at the entry level, surfacing cases whose intermediate states match the active case and returning the parent-case t

  15. Sep 17, 2026 · UTC · The Hacker News

    ThreatsDay: Self-Rewriting Agents, 800+ Flaws Patched, Insider SIM Swaps and 22 More New Stories

    Attackers keep finding new keys. The funny part is that defenders keep inventing where to store them. This week, those keys sit in AI tools, exposed services, old bugs, weak logins, and software sold like a monthly subscription. Some attacks use new tricks. Others just reuse what was already lying around. Both work often enough. So the threat landscape is not getting cleaner. It is just

  16. Sep 17, 2026 · UTC · arXiv · Artificial Intelligence

    Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

    Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters

  17. Sep 17, 2026 · UTC · arXiv · Artificial Intelligence

    Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

    Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repeats. Record-and-replay makes a run reproducible, but existing agent tooling records runs only to trace or score them, not to test a code change against them. We present Chronicle, which records an agent run at its non-deterministic boundaries as immutable envelopes and replays it from the record. Its central operation, cut-point replay, se

  18. Sep 17, 2026 · UTC · AWS Artificial Intelligence Blog

    A shared agentic platform for Wood Mackenzie, on Amazon Bedrock AgentCore

    Wood Mackenzie built APEX, a shared agentic AI platform on Amazon Bedrock AgentCore so every team can ship production agents without rebuilding runtime, identity, observability, and guardrails from scratch. Learn why they chose AgentCore, how APEX Studio operates it, and where multi-agent systems go next.

  19. Sep 17, 2026 · UTC · AWS Artificial Intelligence Blog

    How MRH Trowe enabled secure self-service AI agents in financial services

    Learn how MRH Trowe, one of Germany's leading commercial and industrial insurance brokers, gave about 400 employees secure, self-service access to AI agents in its first month of production - using Strands Agents, Amazon Bedrock AgentCore, and LibreChat to meet the security, data residency, and compliance requirements of the German financial sector.

  20. Sep 17, 2026 · UTC · arXiv · Artificial Intelligence

    Mitigating Retaliatory Algorithmic Collusion in Repeated Games

    Reinforcement learning agents trained to maximize their own reward in repeated interactions can converge to supra-competitive outcomes resembling explicit collusion, without communication or shared design. Existing mitigation approaches are largely tied to specific economic settings, like two-sided platforms and auctions, leaving open how to design interventions for general repeated games. We address this gap by formalizing the connection between empirical observations from prior work on Q-learning collusion and classical theory of Simple Penal Codes (SPCs). We show any non-trivial SPC induces

Explore full timeline