{"id":22992,"date":"2017-04-11T19:26:34","date_gmt":"2017-04-11T19:26:34","guid":{"rendered":"https:\/\/scannn.com\/scanning-7-6-petabytes-of-huggingface-training-data-for-secrets-%e2%97%86-truffle-security-co\/"},"modified":"2017-04-11T19:26:34","modified_gmt":"2017-04-11T19:26:34","slug":"scanning-7-6-petabytes-of-huggingface-training-data-for-secrets-%e2%97%86-truffle-security-co","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/scanning-7-6-petabytes-of-huggingface-training-data-for-secrets-%e2%97%86-truffle-security-co\/","title":{"rendered":"Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co."},"content":{"rendered":"\n<div data-framer-component-type=\"RichTextContainer\" style=\"transform:none\">\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\"><strong class=\"framer-text\">tl;dr<\/strong> We scanned every public dataset on Hugging Face, which is where most open AI training data lives. That came to <strong class=\"framer-text\">7.6 petabytes across 187 million files<\/strong>, the largest secret scan of AI training data we know of. We found <strong class=\"framer-text\">221,303 live, unique credentials<\/strong> sitting in 6,003 datasets.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">One of the highest-impact secrets we found had access to <strong class=\"framer-text\">393 GB<\/strong> of PII covering what we estimate to be roughly <strong class=\"framer-text\">3.7% of the global population<\/strong>. More on this will come in a dedicated follow-up. The rest of the scan shows how broad the problem is: cloud storage buckets, hosted databases, cloud-admin keys, and tokens that can push code into software a lot of people install.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">We shared the findings with Hugging Face before publication; the company partnered closely with us, and CTO Julien Chaumond contributed native storage-bucket scanning support to TruffleHog.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">You\u2019ve probably seen the OpenAI and Hugging Face news. This scan began before that broke, but it\u2019s worth pointing out that part of that kill chain involved stolen API keys. There has never been a stronger imperative for us to work with vendors to get their exposed keys revoked (please reach out to us if we\u2019re not already working with you).<\/p>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">Tokens that can push code into things you install<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">Sometimes when we publish our findings of large quantities of keys, people ask how many of them actually materially matter. Here\u2019s a bunch we found in this scan that have supply chain risk.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">The scariest credentials here let you change software that other people run. Inside the training data we found <strong class=\"framer-text\">349 live GitHub personal access tokens<\/strong>: 223 with full <code class=\"framer-text framer-styles-preset-3yycod\">repo<\/code> write, 130 that can rewrite CI workflows, 112 with <code class=\"framer-text framer-styles-preset-3yycod\">admin:org<\/code>, and 110 that can publish packages. On top of that, <strong class=\"framer-text\">318 Docker Hub tokens that can push images.<\/strong> We checked npm and PyPI specifically and found zero live, so we\u2019re not claiming those.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">A single <code class=\"framer-text framer-styles-preset-3yycod\">repo<\/code> or <code class=\"framer-text framer-styles-preset-3yycod\">admin:org<\/code> token rewrites every repository its owner can push to, and that change ships to everyone who installs the result. Some of these tokens sit on accounts wired into software that millions of people run. Others belonged to accounts positioned deep in the software supply chain.<\/p>\n<div class=\"framer-text framer-text-module\" style=\"width:100%;height:auto\" data-width=\"fill\">\n<div class=\"hfb\">\n<p><span class=\"hfb__leg\"><span class=\"hfb__chip\" style=\"background:#172e29\"\/>GitHub PAT<\/span><span class=\"hfb__leg\"><span class=\"hfb__chip\" style=\"background:#2e86ab\"\/>Docker<\/span><span class=\"hfb__leg\"><span class=\"hfb__chip\" style=\"background:#ee9600\"\/>Hugging Face<\/span><\/p>\n<p><span class=\"hfb__name\">Docker Hub: push images<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:100%;background:#2e86ab;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#2e86ab\">318<\/span><\/p>\n<p><span class=\"hfb__name\">Hugging Face: write<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:74.52830188679245%;background:#ee9600;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#ee9600\">237<\/span><\/p>\n<p><span class=\"hfb__name\">GitHub: full repo write<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:70.12578616352201%;background:#172e29;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#172e29\">223<\/span><\/p>\n<p><span class=\"hfb__name\">GitHub: rewrite CI (workflow)<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:40.88050314465409%;background:#172e29;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#172e29\">130<\/span><\/p>\n<p><span class=\"hfb__name\">GitHub: admin:org<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:35.22012578616352%;background:#172e29;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#172e29\">112<\/span><\/p>\n<p><span class=\"hfb__name\">GitHub: publish packages<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:34.59119496855346%;background:#172e29;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#172e29\">110<\/span><\/p>\n<p><span class=\"hfb__name\">Hugging Face: org-admin<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:22.0125786163522%;background:#ee9600;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#ee9600\">70<\/span><\/p>\n<p>Live, verified tokens \u2014 hover a bar<\/p>\n<\/div>\n<\/div>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\"><em class=\"framer-text\">Live, verified write-capable credentials found in public training data, counted by what they actually authorize.<\/em><\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">One live <code class=\"framer-text framer-styles-preset-3yycod\">repo<\/code>-scoped token belonged to the founder of a widely used Model Context Protocol registry whose account was connected to the official MCP organization. That organization\u2019s repositories hold servers and SDKs used by major AI coding tools and have more than <strong class=\"framer-text\">178,000 GitHub stars<\/strong> between them. Other examples included a highly privileged token held by an engineer at a large technology company, a developer at a bank, and a researcher at an AI lab. We are withholding the names of the people and organizations involved, and have responsibly disclosed our findings.<\/p>\n<blockquote class=\"framer-text framer-styles-preset-1ofny2i\">\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">Following Julien\u2019s contribution to scan storage buckets, we\u2019ve already scanned and found a vast quantity of new keys we\u2019ll do a follow-up post about.<\/p>\n<\/blockquote>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">Keys with real blast radius<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">The scan also turned up live keys that open real infrastructure: cloud accounts, hosted databases, storage buckets, and messaging platforms. We used these credentials only for verification and metadata-only impact checks, meaning database size stats, Redis memory counters, and CloudWatch S3 bucket-size metrics. We did not read database rows, list object keys, download files, or modify anything. Here is what they unlock.<\/p>\n<div class=\"framer-text framer-text-module\" style=\"width:100%;height:auto\" data-width=\"fill\">\n<div class=\"hfc\">\n<div class=\"hfc__grid\">\n<p><span class=\"hfc__label\">Cloud takeover<\/span><strong class=\"hfc__value\">8,557<\/strong><span class=\"hfc__unit\">GCP service-account keys<\/span><span class=\"hfc__note\">Across 3,811 projects<\/span><\/p>\n<p><span class=\"hfc__label\">Private storage<\/span><strong class=\"hfc__value\">51.7 TB<\/strong><span class=\"hfc__unit\">in non-public S3 buckets<\/span><span class=\"hfc__note\">Confirmed from bucket metadata<\/span><\/p>\n<p><span class=\"hfc__label\">Live databases<\/span><strong class=\"hfc__value\">8,594<\/strong><span class=\"hfc__unit\">working database logins<\/span><span class=\"hfc__note\">3.5 TB measured by metadata<\/span><\/p>\n<p><span class=\"hfc__label\">Impersonation<\/span><strong class=\"hfc__value\">5,885<\/strong><span class=\"hfc__unit\">Slack and Mailgun keys<\/span><span class=\"hfc__note\">Many tied to named workspaces or domains<\/span><\/p>\n<p><span class=\"hfc__label\">Chatbot spread<\/span><strong class=\"hfc__value\">18\u00d7<\/strong><span class=\"hfc__unit\">copies of one pasted AWS key<\/span><span class=\"hfc__note\">Captured once, then mirrored<\/span><\/p>\n<\/div>\n<\/div>\n<\/div>\n<h3 dir=\"auto\" class=\"framer-text framer-styles-preset-1kz3lut\"><strong class=\"framer-text\">Cloud takeover \u00b7 GCP \u2014 8,557 live Google service-account keys, across 3,811 projects<\/strong><\/h3>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">A service-account key is a non-interactive credential for a cloud project. Of the verified examples, 1,926 were Firebase admin keys with database access, one carried the explicit <code class=\"framer-text framer-styles-preset-3yycod\">Owner<\/code> role, and one was a Kubernetes <code class=\"framer-text framer-styles-preset-3yycod\">cluster-admin<\/code>. Project metadata indicated that some were associated with healthcare and payment applications. We are withholding project names and account identifiers.<\/p>\n<h3 dir=\"auto\" class=\"framer-text framer-styles-preset-1kz3lut\"><strong class=\"framer-text\">Cloud storage \u00b7 AWS S3 \u2014 51.7 TB in buckets with public access blocked<\/strong><\/h3>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">The full S3 StandardStorage lower bound was 185 TB, but raw byte count is not enough: S3 can hold public assets, logs, backups, or almost anything. So we checked only bucket-level metadata for the largest accounts. Bucket policy and public-access-block settings confirmed 51.7 TB in buckets configured to block public access. Bucket-name tokens pointed at <code class=\"framer-text framer-styles-preset-3yycod\">prod<\/code>, <code class=\"framer-text framer-styles-preset-3yycod\">backup<\/code>, <code class=\"framer-text framer-styles-preset-3yycod\">cloudtrail<\/code>, <code class=\"framer-text framer-styles-preset-3yycod\">invoice<\/code>, <code class=\"framer-text framer-styles-preset-3yycod\">customer<\/code>, <code class=\"framer-text framer-styles-preset-3yycod\">billing<\/code>, <code class=\"framer-text framer-styles-preset-3yycod\">rds<\/code>, <code class=\"framer-text framer-styles-preset-3yycod\">mongo<\/code>, and <code class=\"framer-text framer-styles-preset-3yycod\">terraform<\/code>. We did not list object keys or read object contents.<\/p>\n<ul dir=\"auto\" class=\"framer-text\">\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\">AWS keys passing STS identity checks \u2014 3,343<\/p>\n<\/li>\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\">Keys able to list S3 buckets \u2014 907<\/p>\n<\/li>\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\">Bucket count visible through metadata \u2014 8,676<\/p>\n<\/li>\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\">Buckets with all public-access-block flags enabled \u2014 51.7 TB; largest measured account \u2014 66.9 TB<\/p>\n<\/li>\n<\/ul>\n<div class=\"framer-text framer-text-module\" style=\"width:100%;height:auto\" data-width=\"fill\">\n<div class=\"hfb\">\n<p><span class=\"hfb__name\">S3 lower bound (all buckets)<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:100%;background:#c48d3f;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#c48d3f\">185 TB<\/span><\/p>\n<p><span class=\"hfb__name\">Confirmed non-public S3<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:27.94594594594595%;background:#e77543;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#e77543\">51.7 TB<\/span><\/p>\n<p><span class=\"hfb__name\">Live databases<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:1.891891891891892%;background:#172e29;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#172e29\">3.5 TB<\/span><\/p>\n<p>Storage reachable by leaked keys \u2014 TB, from size metadata only<\/p>\n<\/div>\n<\/div>\n<h3 dir=\"auto\" class=\"framer-text framer-styles-preset-1kz3lut\"><strong class=\"framer-text\">Live databases \u2014 8,594 verified-live database logins, 3.5 TB by metadata<\/strong><\/h3>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">Connection strings that still authenticate to hosted databases. The target names lean heavily toward tutorials and side projects (<code class=\"framer-text framer-styles-preset-3yycod\">test<\/code>, <code class=\"framer-text framer-styles-preset-3yycod\">myfirstdatabase<\/code>, todo apps), and the median MongoDB cluster was only 2.8 MB. But the tail is real: 89 MongoDB clusters and 5 Postgres databases exceeded 1 GB, and the largest MongoDB cluster exposed 617.7 GB by database-size metadata alone. 6,121 of 6,802 MongoDB credentials still connected; the tail included a SQL Server tied to a <strong class=\"framer-text\">US defense contractor<\/strong> and Postgres sets tied to a <strong class=\"framer-text\">Brazilian federal agency<\/strong>.<\/p>\n<div class=\"framer-text framer-text-module\" style=\"width:100%;height:auto\" data-width=\"fill\">\n<div class=\"hfss\">\n<div class=\"hfss__row\">\n<p><strong class=\"hfss__num\">617.7 GB<\/strong><span class=\"hfss__lbl\">Largest single exposed database, by size metadata<\/span><\/p>\n<p><strong class=\"hfss__num\">6,121 \/ 6,802<\/strong><span class=\"hfss__lbl\">MongoDB logins that still authenticate<\/span><\/p>\n<p><strong class=\"hfss__num\">94<\/strong><span class=\"hfss__lbl\">MongoDB &amp; Postgres clusters over 1 GB<\/span><\/p>\n<\/div>\n<\/div>\n<\/div>\n<h3 dir=\"auto\" class=\"framer-text framer-styles-preset-1kz3lut\"><strong class=\"framer-text\">Impersonation \u00b7 Comms \u2014 5,885 live Slack tokens and Mailgun keys<\/strong><\/h3>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">We found 231 Slack tokens, 99.6% of which identified the associated workspace, plus 5,654 Mailgun keys with 2,470 tied to a custom sending domain. The examples included a Fortune 500 technology workspace and sending domains associated with or resembling major technology and consumer brands. We are withholding the workspace and domain names.<\/p>\n<h3 dir=\"auto\" class=\"framer-text framer-styles-preset-1kz3lut\"><strong class=\"framer-text\">A new leak path \u00b7 Chatbots \u2014 18\u00d7 mirrors from one key pasted into a chatbot<\/strong><\/h3>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">A live AWS key tied to a Brazilian lending fintech reached the training data because someone pasted their <code class=\"framer-text framer-styles-preset-3yycod\">boto3<\/code> code into a chatbot. The conversation was captured by LMSYS-Chat-1M and mirrored about 18 times. We are withholding the company name.<\/p>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">What an attacker walks away with<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">Live, verified credentials grouped by what they control. Each credential type sits in one bucket, counted once.<\/p>\n<div class=\"framer-text framer-text-module\" style=\"width:100%;height:auto\" data-width=\"fill\">\n<div class=\"hfb\">\n<p><span class=\"hfb__name\">Email &amp; messaging<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:100%;background:#b44441;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#b44441\">14.5k<\/span><\/p>\n<p><span class=\"hfb__name\">Cloud infrastructure<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:90.30727985151577%;background:#e77543;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#e77543\">13.1k<\/span><\/p>\n<p><span class=\"hfb__name\">AI provider accounts<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:73.61655324121811%;background:#172e29;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#172e29\">10.7k<\/span><\/p>\n<p><span class=\"hfb__name\">Hosted databases<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:59.15996425379804%;background:#4c934f;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#4c934f\">8.6k<\/span><\/p>\n<p>Live, verified credentials \u2014 hover a bar for detail<\/p>\n<\/div>\n<\/div>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">The risk, quantified: at least $920,000 a year in stolen AI inference<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">The training data is full of keys to the AI providers themselves: <strong class=\"framer-text\">11,496 live<\/strong> across 1,210 datasets, covering OpenAI, Azure OpenAI, Anthropic, Gemini, Groq, and more. We never used any of them, so I can\u2019t tell you the real balances, but I don\u2019t need to. Every provider publishes a default spend cap, and a verified-live key sits on an account with at least that cap.<\/p>\n<div class=\"framer-text framer-text-module\" style=\"width:100%;height:auto\" data-width=\"fill\">\n<div class=\"hfm\">\n<div class=\"hfm__equation\" role=\"img\" aria-label=\"742 OpenAI keys plus 26 Anthropic keys, times a 100 dollar default monthly cap, equals a 76,800 dollar monthly floor\">\n<p><strong class=\"hfm__num\">742 + 26<\/strong><span class=\"hfm__lbl\">OpenAI + Anthropic keys<\/span><\/p>\n<p><span class=\"hfm__op\">\u00d7<\/span><\/p>\n<p><strong class=\"hfm__num\">$100<\/strong><span class=\"hfm__lbl\">entry-tier monthly cap<\/span><\/p>\n<p><span class=\"hfm__op\">=<\/span><\/p>\n<p><strong class=\"hfm__num\">$76,800\/mo<\/strong><span class=\"hfm__lbl\">conservative exposure floor<\/span><\/p>\n<\/div>\n<p class=\"hfm__caveat\">This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used.<\/p>\n<p><strong>Up to $200K\/mo<\/strong><span>Anthropic\u2019s published cap for a single top build-tier account\u2014more than the entire scan cost.<\/span><\/p>\n<\/div>\n<\/div>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">A new OpenAI account that has entered billing gets a <strong class=\"framer-text\">$100\/month<\/strong> usage limit by default. Anthropic\u2019s entry tier is the same, $100 a month. We found <strong class=\"framer-text\">742 live OpenAI keys and 26 live Anthropic keys.<\/strong> Drain each one to just its default cap and that\u2019s <strong class=\"framer-text\">$76,800 a month<\/strong>, about <strong class=\"framer-text\">$920,000 a year<\/strong>, of inference billed to people who have no idea their key is in a dataset. That\u2019s the floor, the number you get if every account is stuck on the lowest tier.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">And it climbs fast. OpenAI\u2019s caps run $100, $500, $1,000, $5,000, then $50,000 a month at the top tier. Anthropic\u2019s top build tier caps at $200,000 a month. We found 160 organization-owned OpenAI keys and 34 machine service-account keys, which are exactly the credentials that sit on funded, high-tier billing and almost never get rotated. A single top-tier key drained to its cap bills more in one month than this entire 7.6-petabyte scan cost to run.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">All of that is before Gemini (1,429 keys), DeepSeek (667), xAI\u2019s Grok (162, with no free tier at all), and 174 enterprise Azure OpenAI deployments, several of them baked into AllenAI\u2019s Dolma 3. Every one of these keys is a live invoice pointed at its owner, and a free seat at a frontier model for whoever finds it first.<\/p>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">The biggest scan we\u2019ve ever run<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">We cloned the public dataset hub end to end: every repository, every branch, every large-file object. Then we flattened Parquet, Arrow, JSONL, archives, and binaries into scannable text and ran TruffleHog with verification on. It worked out to <strong class=\"framer-text\">186.9 million unique files<\/strong> and about <strong class=\"framer-text\">7.6 petabytes<\/strong> of content across roughly <strong class=\"framer-text\">815,000 dataset repositories<\/strong>. About 670,000 of those finished cleanly. The largest web-scale scan we\u2019d done before this topped out around 400 terabytes. This one was about nineteen times bigger.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">The size is only half the story. These are the training sets behind models people actually use. The worst-hit ones are named, card-documented pretraining corpora that open models were built on. We verified every credential we cite against its provider, so they were live when we looked.<\/p>\n<ul dir=\"auto\" class=\"framer-text\">\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\"><strong class=\"framer-text\">186.9M<\/strong> unique files scanned<\/p>\n<\/li>\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\"><strong class=\"framer-text\">7.6 PB<\/strong> of AI training data<\/p>\n<\/li>\n<\/ul>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">How big is 7.6 petabytes?<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\"><strong class=\"framer-text\">7.6 petabytes<\/strong> is hard to picture. A single DVD holds 4.7 GB, so this scan would fill about <strong class=\"framer-text\">1.6 million of them<\/strong>. Stacked into a tower, those discs would stand roughly <strong class=\"framer-text\">1.9 kilometers<\/strong> tall, about as high as <strong class=\"framer-text\">4.4 Empire State Buildings<\/strong> on top of each other.<\/p>\n<div class=\"framer-text framer-text-module\" style=\"width:100%;height:auto\" data-width=\"fill\">\n<div class=\"hfd\">\n<p><svg class=\"hfd__svg\" viewbox=\"75 32 320 536\" role=\"img\" aria-label=\"A tower of 1.6 million DVDs holding 7.6 petabytes reaches about 1.9 km, as tall as 4.4 Empire State Buildings stacked on top of each other.\">\n  <defs>\n    <lineargradient id=\"dvdGrad\" x1=\"0\" y1=\"0\" x2=\"0\" y2=\"1\">\n      <stop offset=\"0\" stop-color=\"#f0a06a\"\/>\n      <stop offset=\"1\" stop-color=\"#e4572e\"\/>\n    <\/lineargradient>\n  <\/defs><\/p>\n<p>  <!-- ground -->\n  <line x1=\"95\" y1=\"520\" x2=\"380\" y2=\"520\" stroke=\"#2a1610\" stroke-width=\"2\"\/>\n<p>  <!-- DVD tower: the 7.6 PB scan stored on discs -->\n  <rect x=\"140\" y=\"70\" width=\"58\" height=\"450\" fill=\"url(#dvdGrad)\" stroke=\"#b8431f\" stroke-width=\"1\"\/>\n  <g stroke=\"#ffffff\" stroke-opacity=\"0.25\" stroke-width=\"1\">\n    <line x1=\"140\" y1=\"120\" x2=\"198\" y2=\"120\"\/><line x1=\"140\" y1=\"180\" x2=\"198\" y2=\"180\"\/>\n    <line x1=\"140\" y1=\"240\" x2=\"198\" y2=\"240\"\/><line x1=\"140\" y1=\"300\" x2=\"198\" y2=\"300\"\/>\n    <line x1=\"140\" y1=\"360\" x2=\"198\" y2=\"360\"\/><line x1=\"140\" y1=\"420\" x2=\"198\" y2=\"420\"\/>\n    <line x1=\"140\" y1=\"470\" x2=\"198\" y2=\"470\"\/>\n  <\/g>\n  <text x=\"169\" y=\"52\" text-anchor=\"middle\" font-family=\"Outfit, sans-serif\" font-size=\"15\" font-weight=\"800\" fill=\"#2a1610\">7.6 PB on DVDs<\/text>\n  <text x=\"169\" y=\"66\" text-anchor=\"middle\" font-family=\"IBM Plex Sans, sans-serif\" font-size=\"11\" fill=\"#5b5048\">1.6M discs<\/text>\n  <line x1=\"116\" y1=\"70\" x2=\"116\" y2=\"520\" stroke=\"#2a1610\" stroke-width=\"1\"\/>\n  <line x1=\"112\" y1=\"70\" x2=\"120\" y2=\"70\" stroke=\"#2a1610\" stroke-width=\"1\"\/>\n  <line x1=\"112\" y1=\"520\" x2=\"120\" y2=\"520\" stroke=\"#2a1610\" stroke-width=\"1\"\/>\n  <text x=\"104\" y=\"298\" text-anchor=\"middle\" font-family=\"Outfit, sans-serif\" font-size=\"15\" font-weight=\"700\" fill=\"#2a1610\" transform=\"rotate(-90 104 298)\">\u2248 1.9 km<\/text><\/p>\n<p>  <!-- equal-height line: DVD tower top to ESB stack top -->\n  <line x1=\"198\" y1=\"70\" x2=\"283\" y2=\"70\" stroke=\"#8a8076\" stroke-width=\"1\" stroke-dasharray=\"4 3\"\/>\n  <text x=\"240\" y=\"62\" text-anchor=\"middle\" font-family=\"IBM Plex Sans, sans-serif\" font-size=\"10\" fill=\"#5b5048\">same height<\/text><\/p>\n<p>  <!-- Empire State Buildings stacked 4.4x (slender: base 38 \/ shaft 26 \/ spire 3, 103 px = 443 m) -->\n  <g fill=\"#7d746b\" stroke=\"#5b5048\" stroke-width=\"1\">\n    <rect x=\"283\" y=\"502\" width=\"38\" height=\"18\"\/><rect x=\"289\" y=\"441\" width=\"26\" height=\"61\"\/><rect x=\"295\" y=\"431\" width=\"14\" height=\"10\"\/><rect x=\"300.5\" y=\"417\" width=\"3\" height=\"14\"\/>\n    <rect x=\"283\" y=\"399\" width=\"38\" height=\"18\"\/><rect x=\"289\" y=\"338\" width=\"26\" height=\"61\"\/><rect x=\"295\" y=\"328\" width=\"14\" height=\"10\"\/><rect x=\"300.5\" y=\"314\" width=\"3\" height=\"14\"\/>\n    <rect x=\"283\" y=\"296\" width=\"38\" height=\"18\"\/><rect x=\"289\" y=\"235\" width=\"26\" height=\"61\"\/><rect x=\"295\" y=\"225\" width=\"14\" height=\"10\"\/><rect x=\"300.5\" y=\"211\" width=\"3\" height=\"14\"\/>\n    <rect x=\"283\" y=\"193\" width=\"38\" height=\"18\"\/><rect x=\"289\" y=\"132\" width=\"26\" height=\"61\"\/><rect x=\"295\" y=\"122\" width=\"14\" height=\"10\"\/><rect x=\"300.5\" y=\"108\" width=\"3\" height=\"14\"\/>\n    <rect x=\"283\" y=\"90\" width=\"38\" height=\"18\"\/><rect x=\"289\" y=\"70\" width=\"26\" height=\"20\"\/>\n  <\/g>\n  <g stroke=\"#ffffff\" stroke-opacity=\"0.55\" stroke-width=\"1\">\n    <line x1=\"283\" y1=\"417\" x2=\"321\" y2=\"417\"\/><line x1=\"283\" y1=\"314\" x2=\"321\" y2=\"314\"\/><line x1=\"283\" y1=\"211\" x2=\"321\" y2=\"211\"\/><line x1=\"283\" y1=\"108\" x2=\"321\" y2=\"108\"\/>\n  <\/g>\n  <text x=\"338\" y=\"298\" text-anchor=\"start\" font-family=\"Outfit, sans-serif\" font-size=\"22\" font-weight=\"800\" fill=\"#2a1610\">4.4\u00d7<\/text>\n  <text x=\"302\" y=\"543\" text-anchor=\"middle\" font-family=\"Outfit, sans-serif\" font-size=\"14\" font-weight=\"700\" fill=\"#2a1610\">Empire State Buildings<\/text>\n  <text x=\"302\" y=\"560\" text-anchor=\"middle\" font-family=\"IBM Plex Sans, sans-serif\" font-size=\"12\" fill=\"#5b5048\">443 m each, stacked<\/text>\n<\/svg><\/p>\n<\/div>\n<\/div>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">Sometimes the company leaked its own key<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">There are two ways a live key ends up in a dataset. Most of the time it belongs to a stranger. It leaked somewhere public, got scraped, and rode a corpus into a dataset whose publisher has never heard of them. The rarer case is also the more awkward one: the people who built and published the dataset leaked their own working key into it. Usually it\u2019s the exact token they use to push to Hugging Face, sitting in a notebook, a cache file, or a backup they uploaded by accident.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">The training data holds <strong class=\"framer-text\">787 live Hugging Face tokens<\/strong>, 237 with <strong class=\"framer-text\">write<\/strong> access and 70 with <strong class=\"framer-text\">org-admin<\/strong>. A live write token in a public dataset is a key to push malicious model weights or poison datasets across a whole org. Of the ones we could trace, about <strong class=\"framer-text\">700 had been scraped into someone else\u2019s corpus<\/strong> and <strong class=\"framer-text\">63 were self-leaks<\/strong>, dropped by the owner into their own namespace. Either way, the AI supply chain is leaking the exact keys an attacker would need to poison it, and a tampered model or dataset can ride the same pipeline straight to everyone downstream.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">Hugging Face already kills its own tokens in the most direct case: for Enterprise organizations, an HF token pushed into a public repo or bucket is auto-revoked on the spot. The live Hub tokens we found are the ones that fall outside that net, meaning self-leaks on non-Enterprise accounts and tokens scraped in from someone else\u2019s corpus.<\/p>\n<div class=\"framer-text framer-text-module\" style=\"width:100%;height:auto\" data-width=\"fill\">\n<div class=\"hfb\">\n<p><span class=\"hfb__name\">Scraped from a stranger<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:100%;background:#2e86ab;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#2e86ab\">700<\/span><\/p>\n<p><span class=\"hfb__name\">Self-leaked by the owner<\/span><span class=\"hfb__track\"><span class=\"hfb__fill\" style=\"width:9%;background:#b44441;opacity:1;filter:none\"\/><\/span><span class=\"hfb__val\" style=\"color:#b44441\">63<\/span><\/p>\n<p>How 763 traceable Hugging Face write &amp; org-admin tokens reached a dataset<\/p>\n<\/div>\n<\/div>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">Some self-leaks were tied directly to people with publishing access. The head of product at an AI infrastructure startup left an account-write Hub token in the company\u2019s own text-to-SQL benchmark dataset. Another developer committed a highly privileged GitHub token with organization administration, repository write, CI workflow, package publishing, and repository deletion permissions into a speech dataset. An administrator across several pretraining-data organizations exposed a write token in one of those organizations\u2019 datasets. We also found corporate upload tokens in datasets published by a data-labeling vendor and researchers at two AI labs. In each case, the credential was published by someone connected to the account or organization it could modify.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">The cloud keys break the other way. Every live <strong class=\"framer-text\">AWS, GCP, Azure, Docker, and database<\/strong> credential we could attribute traced back to someone other than the publisher. Almost nobody pastes their own live cloud key into a public dataset on purpose, so those got there by being scraped. That\u2019s the next story.<\/p>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">One leak, copied into hundreds of datasets<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">Most of these secrets leaked somewhere else first, in a GitHub repo or a web page or a chat log. From there they got vacuumed into an upstream corpus like The Stack or Common Crawl, and then rode every derivative of that corpus downstream. <strong class=\"framer-text\">44% of all unique live secrets show up in more than one dataset.<\/strong> 19,380 of them appear in ten or more. The Stack and its forks alone carry <strong class=\"framer-text\">51,571 distinct live keys<\/strong>, and AllenAI\u2019s Dolma family carries <strong class=\"framer-text\">28,110<\/strong>. Of the keys that reach ten or more datasets, <strong class=\"framer-text\">99.3% pass through one of these scrape corpora.<\/strong> The amplification is mechanical. Scrape a corpus once, and every dedup, filter, and fine-tuning remix republishes the same live keys.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">The same verified credentials flow from upstream source families into the big pretraining corpora. As an example, two near-identical <code class=\"framer-text framer-styles-preset-3yycod\">stack-edu<\/code> re-uploads share 19,977 of 19,977 keys, and Stack code keys cross straight into AllenAI\u2019s Dolma 3 web mixes.<\/p>\n<div class=\"framer-text framer-text-module\" style=\"width:100%;height:auto\" data-width=\"fill\">\n<div class=\"hfs\"><svg viewbox=\"0 0 920 540\" width=\"100%\" role=\"img\" aria-label=\"Sankey of secret source families flowing into the most-affected datasets\"><g fill=\"none\"><path class=\"hfs__link\" d=\"M22,26.611445902150795 C460,26.611445902150795 460,26.611445902150795 898,26.611445902150795\" stroke=\"#e4572e\" stroke-opacity=\"0.28\" stroke-width=\"25.22289180430159\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,139.61703208470004 C460,139.61703208470004 460,50.27839272797662 898,50.27839272797662\" stroke=\"#2e86ab\" stroke-opacity=\"0.28\" stroke-width=\"22.11100184735006\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,160.04502234302294 C460,160.04502234302294 460,280.7792721749622 898,280.7792721749622\" stroke=\"#2e86ab\" stroke-opacity=\"0.28\" stroke-width=\"18.74497866929574\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,47.04889629095851 C460,47.04889629095851 460,297.977765996267 898,297.977765996267\" stroke=\"#e4572e\" stroke-opacity=\"0.28\" stroke-width=\"15.652008973313844\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,62.06646653596141 C460,62.06646653596141 460,380.45016154533175 898,380.45016154533175\" stroke=\"#e4572e\" stroke-opacity=\"0.28\" stroke-width=\"14.383131516691941\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,177.31983641738708 C460,177.31983641738708 460,395.544052043394 898,395.544052043394\" stroke=\"#2e86ab\" stroke-opacity=\"0.28\" stroke-width=\"15.80464947943257\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,76.14372022589455 C460,76.14372022589455 460,424.2676916456424 898,424.2676916456424\" stroke=\"#e4572e\" stroke-opacity=\"0.28\" stroke-width=\"13.77137586317434\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,192.8268801890039 C460,192.8268801890039 460,438.7580986091301 898,438.7580986091301\" stroke=\"#2e86ab\" stroke-opacity=\"0.28\" stroke-width=\"15.209438063801011\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,89.89516247788684 C460,89.89516247788684 460,466.3193567664294 898,466.3193567664294\" stroke=\"#e4572e\" stroke-opacity=\"0.28\" stroke-width=\"13.73150864081027\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,208.01427338555928 C460,208.01427338555928 460,480.76778525148944 898,480.76778525148944\" stroke=\"#2e86ab\" stroke-opacity=\"0.28\" stroke-width=\"15.165348329309785\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,103.04550324785541 C460,103.04550324785541 460,328.9177015665186 898,328.9177015665186\" stroke=\"#e4572e\" stroke-opacity=\"0.28\" stroke-width=\"12.569172899126826\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,114.44581042922191 C460,114.44581042922191 460,506.5451818186332 898,506.5451818186332\" stroke=\"#e4572e\" stroke-opacity=\"0.28\" stroke-width=\"10.231441463606192\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,226.39247107395818 C460,226.39247107395818 460,126.02954652087406 898,126.02954652087406\" stroke=\"#2e86ab\" stroke-opacity=\"0.28\" stroke-width=\"21.59104704748801\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,261.8508849289611 C460,261.8508849289611 460,147.21923710417542 898,147.21923710417542\" stroke=\"#8338ec\" stroke-opacity=\"0.28\" stroke-width=\"20.78833411911472\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,301.0642433364037 C460,301.0642433364037 460,224.13527995999152 898,224.13527995999152\" stroke=\"#8338ec\" stroke-opacity=\"0.28\" stroke-width=\"57.638382695770424\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,349.29467719637034 C460,349.29467719637034 460,71.7451361637331 898,71.7451361637331\" stroke=\"#6c757d\" stroke-opacity=\"0.28\" stroke-width=\"20.822485024162887\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><path class=\"hfs__link\" d=\"M22,239.82235623355297 C460,239.82235623355297 460,337.8366496519328 898,337.8366496519328\" stroke=\"#2e86ab\" stroke-opacity=\"0.28\" stroke-width=\"5.26872327170156\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/path><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"8\" y=\"14\" width=\"14\" height=\"105.561531161025\" fill=\"#e4572e\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"28\" y=\"66.78076558051251\" dy=\"0.35em\" text-anchor=\"start\" font-weight=\"700\" fill=\"#e4572e\">The Stack<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"8\" y=\"128.56153116102502\" width=\"14\" height=\"113.89518670837873\" fill=\"#2e86ab\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"28\" y=\"185.50912451521438\" dy=\"0.35em\" text-anchor=\"start\" font-weight=\"700\" fill=\"#2e86ab\">Common Crawl \/ Web<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"8\" y=\"251.45671786940375\" width=\"14\" height=\"78.42671681488514\" fill=\"#8338ec\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"28\" y=\"290.6700762768463\" dy=\"0.35em\" text-anchor=\"start\" font-weight=\"700\" fill=\"#8338ec\">Git commits \/ events<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"8\" y=\"338.8834346842889\" width=\"14\" height=\"187.1165653157111\" fill=\"#6c757d\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"28\" y=\"432.44171734214444\" dy=\"0.35em\" text-anchor=\"start\" font-weight=\"700\" fill=\"#6c757d\">Other<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"898\" y=\"14\" width=\"14\" height=\"92.23402299713005\" fill=\"#6c757d\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"892\" y=\"60.11701149856503\" dy=\"0.35em\" text-anchor=\"end\" font-weight=\"600\" fill=\"#2a1610\">huginn-dataset<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"898\" y=\"115.23402299713005\" width=\"14\" height=\"71.08206561497626\" fill=\"#2e86ab\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"892\" y=\"150.77505580461818\" dy=\"0.35em\" text-anchor=\"end\" font-weight=\"600\" fill=\"#2a1610\">latent-cot-nemotron<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"898\" y=\"195.3160886121063\" width=\"14\" height=\"67.09069422820804\" fill=\"#8338ec\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"892\" y=\"228.86143572621035\" dy=\"0.35em\" text-anchor=\"end\" font-weight=\"600\" fill=\"#2a1610\">diffs<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"898\" y=\"271.40678284031435\" width=\"14\" height=\"42.226332276640846\" fill=\"#2e86ab\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"892\" y=\"292.5199489786348\" dy=\"0.35em\" text-anchor=\"end\" font-weight=\"600\" fill=\"#2a1610\">dolma3_dolmino_pool<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"898\" y=\"322.6331151169552\" width=\"14\" height=\"41.62548067003058\" fill=\"#e4572e\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"892\" y=\"343.4458554519705\" dy=\"0.35em\" text-anchor=\"end\" font-weight=\"600\" fill=\"#2a1610\">Stack_Tokenized<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"898\" y=\"373.2585957869858\" width=\"14\" height=\"35.12340792706948\" fill=\"#2e86ab\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"892\" y=\"390.8202997505205\" dy=\"0.35em\" text-anchor=\"end\" font-weight=\"600\" fill=\"#2a1610\">dolma3_mix-6T<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"898\" y=\"417.38200371405526\" width=\"14\" height=\"33.071598731969004\" fill=\"#e4572e\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"892\" y=\"433.9178030800398\" dy=\"0.35em\" text-anchor=\"end\" font-weight=\"600\" fill=\"#2a1610\">stack-edu<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"898\" y=\"459.45360244602426\" width=\"14\" height=\"32.975858640805825\" fill=\"#e4572e\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"892\" y=\"475.94153176642715\" dy=\"0.35em\" text-anchor=\"end\" font-weight=\"600\" fill=\"#2a1610\">stack-edu<\/text><\/g><g class=\"hfs__node\" style=\"opacity:1\"><rect x=\"898\" y=\"501.4294610868301\" width=\"14\" height=\"24.570538913169912\" fill=\"#e4572e\"><title>Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets \u25c6 Truffle Security Co.<\/title><\/rect><text class=\"hfs__lbl\" x=\"892\" y=\"513.714730543415\" dy=\"0.35em\" text-anchor=\"end\" font-weight=\"600\" fill=\"#2a1610\">stack-edu-python<\/text><\/g><\/svg><\/p>\n<p>Source family \u2192 dataset. Width \u221d shared live secrets. Hover to trace a flow.<\/p>\n<\/div>\n<\/div>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\"><em class=\"framer-text\">Flow weight is the number of identical verified-live secrets (by credential hash) an upstream family shares with a dataset. Showing the four source families and the fifteen most secret-laden derivative datasets.<\/em><\/p>\n<blockquote class=\"framer-text framer-styles-preset-1ofny2i\">\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\"><strong class=\"framer-text\">1,131 datasets contained the same single live key.<\/strong> One Infura key, pasted once into a ChatGPT conversation, got captured by the WildChat chat-log dataset and then copied into 1,131 public datasets and 10,162 file locations. Revoking it at the source does nothing about the other 1,130 copies. A leak in training data spreads on its own, and every copy is another place it keeps working.<\/p>\n<\/blockquote>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">Models you\u2019ve heard of, trained on live keys<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">The household-name chatbots like ChatGPT, Claude, Gemini, and Llama keep their training data private, so I can\u2019t tell you what\u2019s in them. The open models that publish their data are full of live keys.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\"><em class=\"framer-text\">Unique live secrets per dataset, deduplicated by credential. Model attribution confirmed from each dataset\u2019s public card. Downloads are Hugging Face all-time.<\/em><\/p>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">If you touch any part of this pipeline<\/strong><\/h2>\n<h3 dir=\"auto\" class=\"framer-text framer-styles-preset-1kz3lut\"><strong class=\"framer-text\">Dataset authors &amp; AI labs<\/strong><\/h3>\n<ul dir=\"auto\" class=\"framer-text\">\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\"><strong class=\"framer-text\">Scan before you publish.<\/strong> Run a secret scanner over a corpus before it goes up. It\u2019s the cheapest step in the pipeline, and it stops you from shipping live keys into every downstream remix.<\/p>\n<\/li>\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\"><strong class=\"framer-text\">Scan before you train.<\/strong> If you\u2019re pulling a public dataset into a run, assume it has live credentials until you\u2019ve checked. The big code and web corpora clearly do.<\/p>\n<\/li>\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\"><strong class=\"framer-text\">Keep your leak alerts on.<\/strong> If you publish on Hugging Face, it already scans your pushes with TruffleHog and emails you the moment a verified secret lands. Those alerts are on by default, so check you haven\u2019t switched them off. Enterprise orgs also get Hugging Face tokens auto-revoked when they hit a public repo or bucket.<\/p>\n<\/li>\n<\/ul>\n<h3 dir=\"auto\" class=\"framer-text framer-styles-preset-1kz3lut\"><strong class=\"framer-text\">Developers &amp; providers<\/strong><\/h3>\n<ul dir=\"auto\" class=\"framer-text\">\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\"><strong class=\"framer-text\">Rotate, don\u2019t hide.<\/strong> Any key that ever hit a public repo, a web page, or a chatbot should be treated as burned. Rotation is the only fix that survives being copied into a dataset.<\/p>\n<\/li>\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\"><strong class=\"framer-text\">Offer bulk revocation.<\/strong> Providers with revocation APIs and proactive scanning let researchers help at this scale. Without them, 221,303 keys is just a spreadsheet nobody can act on.<\/p>\n<\/li>\n<\/ul>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">Training data is the most permanent leak there is<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">You can scrub a secret out of git history with a force-push. A training set has no undo. By the time a key lands in a published corpus, it\u2019s been copied into derivative datasets, downloaded onto thousands of machines, and folded into model weights. These datasets are valuable because they\u2019re permanent, versioned, and remixed constantly. That is exactly what makes a leaked key in one impossible to clean up.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">The fix hasn\u2019t changed: <strong class=\"framer-text\">rotate the key.<\/strong> Once you revoke it, the copy in the dataset is just a harmless string. Until then it keeps working every time the data gets reused, and reuse is the whole point of training data.<\/p>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">221,303 keys is hard to disclose<\/strong><\/h2>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">We always try to help people revoke what we find, and at this scale we started before publishing. The highest-impact findings, including the exposure behind that 393 GB of PII, went to the affected parties and their providers ahead of this post so they could revoke and lock things down first, and we are holding publication until the most critical of them confirm receipt. We are not naming any of the people, companies, or datasets involved. Emailing 221,303 owners one by one isn\u2019t realistic, and most of them have no idea what a training dataset even is, let alone that they\u2019re in one, so we are also working the provider side: notifying the vendors whose customers are most affected so they can revoke in bulk, and routing verified findings through the partner channels we already have.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">Hugging Face and the dataset authors didn\u2019t cause this. They\u2019re publishing snapshots of public code and the public web, which is what they\u2019re supposed to do. The secrets leaked upstream, from people who\u2019ll probably never read this. Anyone training on this data should still know it\u2019s in there. Nothing here is a how-to: we\u2019ve withheld the dataset names, file paths, and key material that would let a reader pull live credentials out of these corpora.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">Hugging Face isn\u2019t passive about this either: they already run <!--$--><a class=\"framer-text framer-styles-preset-1bc2aw\" href=\"https:\/\/huggingface.co\/docs\/hub\/en\/security-secrets\" rel=\"\">TruffleHog on every push to a public repo and email the author when a verified secret appears<\/a><!--\/$-->, part of <!--$--><a class=\"framer-text framer-styles-preset-1bc2aw\" href=\"https:\/\/huggingface.co\/blog\/trufflesecurity-partnership\" rel=\"\">our ongoing partnership<\/a><!--\/$-->. The catch is the gap this whole scan measures: a notification only helps if someone acts on it, and the rate of action is well short of 100%. Detected, notified, but never revoked is a big part of why so many of these keys were still live when we looked.<\/p>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\">We shared these findings with Hugging Face ahead of publication, and they\u2019ve been a real partner in getting ahead of the problem. Their CTO, <strong class=\"framer-text\">Julien Chaumond<\/strong>, went a step further and contributed code directly to TruffleHog: native scanning support for Hugging Face\u2019s new storage buckets, so secret scans of an org or user now cover object storage alongside models, datasets, and Spaces. That support lands in an upcoming TruffleHog release, and we\u2019ll share more about the collaboration in a future post.<\/p>\n<h2 dir=\"auto\" class=\"framer-text framer-styles-preset-9d3ek\"><strong class=\"framer-text\">Takeaways<\/strong><\/h2>\n<ul dir=\"auto\" class=\"framer-text\">\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\">AI training data is full of live credentials. 221,303 unique, verified keys across 6,003 public datasets, including the named training corpora of models people actually use.<\/p>\n<\/li>\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\">What a key unlocks matters more than its raw permission level. The damage here comes from credentials that push code into installed software, open hosted databases, take over cloud accounts, and send mail as real brands.<\/p>\n<\/li>\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\">Leaks multiply. 44% of these keys live in more than one dataset, and one reached 1,131. Chat logs are now their own leak path, capturing keys people paste into chatbots.<\/p>\n<\/li>\n<li data-preset-tag=\"p\" class=\"framer-text framer-styles-preset-11g5x76\">\n<p class=\"framer-text framer-styles-preset-11g5x76\">Rotation is still the only fix. Scan corpora before publishing, scan before training, and revoke anything that ever touched a public surface.<\/p>\n<\/li>\n<\/ul>\n<p dir=\"auto\" class=\"framer-text framer-styles-preset-11g5x76\"><em class=\"framer-text\">Research and analysis by the Truffle Security research team. Scanning powered by TruffleHog. Live credentials were used only for verification and metadata-only impact checks; no stored data was read, copied, or modified.<\/em><\/p>\n<\/div>\n<p><a href=\"https:\/\/trufflesecurity.com\/blog\/scanning-7-6-petabytes-of-ai-training-data-for-secrets?utm_source=tldrit\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>tl;dr We scanned every public dataset on Hugging Face, which is where most open AI training data lives. That came to 7.6 petabytes across 187 million files, the largest secret scan of AI training data we know of. We found 221,303 live, unique credentials sitting in 6,003 datasets. One of the highest-impact secrets we found [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":22993,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-22992","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22992","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=22992"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22992\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/22993"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=22992"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=22992"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=22992"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}