AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Alleged Inclusion of 12,000 Live API Keys in LLM Training Data Reportedly Poses Security Risks

AI Incident Database · article · Feb 28, 2025 · UTC

A dataset used to train large language models allegedly contained 12,000 live API keys and authentication credentials. Some of these were reportedly still active and allowed unauthorized access. Truffle Security found these secrets in a December 2024 Common Crawl archive, which spans 250 billion web pages. The affected credentials could have been exploited for unauthorized data access, service disruptions, financial fraud, and a variety of other malicious uses.

Read original source ↗ Open in workspace

recordType
incident-report
evidenceStatus
reported
region
Global

Reported occurrence date: 2025-02-28T00:00:00.000Z

Evidence & attribution

AI Incident Database, Responsible AI Collaborative; McGregor (2021), Preventing Repeated Real World AI Failures by Cataloging Incidents. Incident-specific contributor credits are available at each citation link. Metadata adapted; article text excluded.

License: CC BY-SA 4.0

First collected: 2026-09-19T22:50:59.123Z. This is not the publication date.

Observed changes

AIIC observation times, not verified publisher revision times. Up to eight recent revisions.

2026-09-20T23:22:28.549Z

  • publishedAt: Not provided2025-02-28T00:00:00.000Z