The AI safety-evaluation organization METR, the very research body that investigated OpenAI agents attacking Hugging Face, recently disclosed external attacks that took place earlier this year. In one incident, the API credentials for an AI model that METR used were stolen. After this happened, over roughly 20 days, the credentials consumed about $600,000 worth of AI credits.
An Authentication Failure Exposed the Agent Dashboard
A METR researcher ran an agent-orchestration dashboard on a personal server. The panel was originally meant to be protected by a Google account. However, the application itself was AI-built and therefore contained a vulnerability. As detailed in METR’s security update, when an authentication anomaly occurred, the application did not deny access. Instead, it silently disabled its authentication function. Thus, the dashboard remained exposed on the open internet for several days.
METR surmises that the attacker may have searched certificate-transparency logs for newly registered domains. Then, the attacker looked for services bearing keywords such as LLM and agent. Having found a target, the attacker prompted the agent directly, inducing it to reveal the model provider’s API credentials. In addition, the attacker added an SSH key pair to establish persistent access.
$600,000 in Anomalous Usage Went Undetected
The stolen API key could only access public models. METR holds special permissions with numerous AI companies, granting access to unreleased models, but this key did not. The attacker used the credentials to call the relevant models directly. As a result, they ran up roughly $600,000 in charges.
The credentials were most likely sold to a downstream API-reselling service for profit. Otherwise, it is unlikely that $600,000 in credits could have been consumed in a mere 20 days. The credits themselves, however, had been granted to METR for free by another AI company. Therefore, METR did not actually have to pay.
METR says the anomalous usage went unnoticed for so long because the organization routinely consumes enormous volumes of credits during its evaluations. Moreover, at the time, some API keys could not have a spending cap set. Additionally, the model console did not show researchers the full picture of rate-limited requests.
A Sustained Campaign of Malicious Scanning
METR also noted a separate campaign in which attackers relentlessly probed its infrastructure for vulnerabilities. The attempted operations included credential-stuffing attacks, OAuth authorization attempts, scans of newly deployed services, and phishing aimed at METR staff.
METR’s public transcript viewer inadvertently exposed a read-only SQL interface. In theory, the bug in this interface could have accessed unpublished model-evaluation data. Also, the database had mistakenly mixed in a small amount of sensitive model output.
An audit found, however, that the vulnerability was discovered and responsibly disclosed by a white-hat security researcher. Although malicious scanners probed the endpoint in passing, they did not successfully exploit it, nor did they actually access METR’s not-yet-published data.
Support Our Threat Intelligence
Find our zero-day alerts and CVE reports helpful? Support our work today and unlock a 100% ad-free reading experience!