Business · Xiaohu's Take

Hacked Suno Code Exposes Training Library: Scraped Roughly 380,000 Hours From YouTube Music and Others

A worm breached Suno, and leaked source code shows the sites it scraped from and for how long. User data also surfaced; the company says it doesn't need to notify users individually.

The TL;DR
  • A hacker (alias ellie.191) used a worm called Shai-Hulud to compromise Suno employee accounts and handed source code related to training data to 404 Media.
  • Code comments list scraping from YouTube Music, Pond5, Deezer, Genius, and others; nine libraries total about 380,000 hours. One file, youtube_music, records over 2.01 million audio segments scraped.
  • Scraping used Bright Data proxies to pull from YouTube and specifically searched for a cappella tracks. There was also a plan to download roughly 1 million hours from about 420,000 podcasts; this was the intended scope, not necessarily a record of completed downloads.
  • The same breach exposed Suno user emails, phone numbers, and Stripe-related payment info. The hacker told 404 Media it involved several hundred thousand users. TechCrunch reported some card data was also present.
  • Suno says the incident was detected in November 2025, had limited impact, mainly involved superseded code, and that it couldn't access full card numbers, so it wasn't legally obligated to notify users individually. Some users told 404 Media they never received any notification.
1The Setup

Suno gets hacked, revealing two things

AI music company Suno was breached. A hacker known as ellie.191 used a worm called Shai-Hulud (literally, "sandworm") to compromise employee login credentials, obtain source code related to training data, and pass it to tech outlet 404 Media.

The most striking part of the leaked code is the scale and sources of the training set: YouTube Music, Deezer, Genius, and Pond5 are all named in comments, adding up to roughly 380,000 hours. There's also a plan to download about 1 million hours of podcasts.

  • 113,879 hours · YouTube Music (youtube_music)
  • 152,162 hours · Another YouTube library (ytm_tagged, the largest single item)
  • 62,117 hours · Pond5 stock library
  • 12,287 hours · Deezer, plus Genius lyrics data
  • ~1 million hours · Podcast download plan (approx. 420,000 shows)

In the same breach, the hacker also accessed Suno user emails, phone numbers, and payment-related data from Stripe. They told 404 Media it involved "several hundred thousand" users. TechCrunch, in its coverage, noted that partial card numbers were among the accessed material.

Suno's response: the security incident was detected in November 2025, was quickly contained, and had limited impact, so individual notification wasn't required. Some users who spoke to 404 Media, however, said they only learned about the potential exposure from the news.

There are two key threads here. First, the specifics of what Suno scraped and from where — previously vague, now detailed in code comments. Second, whether the company adequately notified users after the breach. 404 Media published first on July 15, 2026; Music Business Worldwide, Variety, and TechCrunch followed.
What is Shai-Hulud?

Shai-Hulud is a supply-chain worm: instead of attacking the company directly, it targets the tools and credentials on developers' machines to steal GitHub, cloud, and other accounts. Hacker ellie.191 told 404 Media they used this worm to compromise a Suno employee. When asked why they targeted Suno, the hacker's response was essentially: they like to hack everything.

Lead image from 404 Media
Lead image from 404 Media. Credit: 404 Media (Image: Suno)
2The Scale

Where the data came from, and how much there is

The source code seen by media outlets dates from roughly 2023 to 2024. One set of comments outlines data sources, noting non-music content will be filtered out.

Libraries named in code comments

genius_hq, youtube_music, freesound, jamendo, imp (IMSLP), deezer, ytm_tagged

Other datasets like pond5_music and musescore_lyrics are also referenced. Comments included the note "non-music will be filtered out."

~382,000
Hours · Total across nine libraries (roughly 43 years of continuous audio)
~266,000
Hours · Combined total for two YouTube-related libraries
2.01M+
Segments · Counted in the youtube_music file
~1 million
Hours · Podcast download target (detailed in the next section)
Dataset sizes by hour
ytm_tagged
152,162
youtube_music
113,879
pond5_music
62,117
imslp
19,514
genius_hq
17,615
deezer
12,287
jamendo
3,726
freesound
410
musescore
103

Longer bars mean more hours. ytm_tagged is the largest. Adding these nine figures gives roughly 381,813 hours total; 404 Media's article doesn't separate out a combined sum, noting only that the total represents "at least decades of music." The ~1 million hours of podcasts are a separate planned line and not included in this chart.

Streaming
YouTube Music
youtube_music and ytm_tagged together total roughly 266,000 hours. The youtube_music file also notes over 2.01 million segments scraped.
Streaming
Deezer
Roughly 12,300 hours. Deezer typically requires a paid subscription for full tracks. Suno claims it trains only on "publicly available" material from the open internet; how both are true isn't detailed in public documents.
Lyrics Site
Genius
Roughly 17,600 hours. Genius primarily hosts lyrics, often embedding full songs from Apple Music rather than hosting the tracks themselves.
Stock Library
Pond5, etc.
Pond5 is owned by Shutterstock and its public materials list around 2.5 million audio tracks. At the noted 62,117 hours, outlets like MBW suggest Suno took a significant portion.
3The Method

How the code went about scraping

So much for how much. Here's how it was done. Music Business Worldwide and other outlets pulled additional specifics from 404 Media's materials.

Target YouTube, etc. Bright Data Proxy scraping Searches a cappella Filter Non-music out Suno Model Training data
Diagram: The YouTube scraping pipeline. 404 Media notes the files don't fully detail how other sites were scraped.

YouTube: Proxies and a cappella searches

The code shows Suno used Bright Data's proxy service to scrape tracks from YouTube. It also searched for "a cappella" — vocal-only versions — seemingly with the goal of isolating singing voices.

The stream ripping that labels accuse Suno of means recording audio while it plays in real time to save it as a file. The process in the source code is clearly batch-run via scripts, which is different from a model occasionally listening to a few tracks online.

Podcasts: A plan for about 1 million hours

There's another line in the code: using PodcastIndex to identify roughly 420,000 podcasts. The filters appear to require at least 5 episodes per show and episodes of roughly 30 minutes, with the goal of downloading about 1 million hours of audio.

That 1 million hours is described as a download plan. Public reporting doesn't confirm it all was completed or used for training. The nine libraries totaling ~380,000 hours are a separate, clearly annotated figure — don't mix them up.

What the company says about "public data"

Suno has stated in court documents and its California disclosure page that training uses publicly accessible music files online and that it respects paywalls and password protection. Yet Deezer and Pond5, which are named, typically require payment or a subscription. How Suno reconciles "respecting paywalls" with obtaining these sources isn't detailed in available public materials.

4The Lawsuits

How this connects to the litigation

Suno is already in court with major record labels. The leaked source code fills in details about "what" and "how" it scraped.

2024 onwards · Major label lawsuits
Under the coordination of the RIAA, Universal, Sony, and others sued Suno for using vast amounts of copyrighted music for training without authorization. Suno argues this is fair use. Suno has acknowledged its training set covers nearly all music files of decent audio quality on the open internet, totaling tens of millions of recordings.
Sept 2025 · Lawsuits amended
The RIAA further alleged Suno engaged in stream ripping from YouTube and circumvented anti-download technical measures (like the rolling cipher). This could trigger the DMCA's anti-circumvention provision. The anti-circumvention claim is a separate track from the fair use defense and can be litigated independently.
Nov 2025 · Company says it found an incident
Suno says it detected a limited security incident at that time and quickly contained it. The broader public discussion of training code and user data came after media reports in July 2026.
Around May 2026 · Damage claims grow
According to Music Business Worldwide, Universal and Sony sought to expand the number of works in question from 560 tracks to around 61,000 (identified via audio fingerprinting). At maximum statutory damages per work, the theoretical ceiling could rise from about $84 million to over $9 billion. The judge has not yet ruled on whether to allow the expansion.
June 29, 2026 · Jamendo sues too
Music licensing platform Jamendo filed its own lawsuit against Suno in the same US court. Jamendo claims about 55,600 tracks were licensed only for non-commercial academic use, and Suno used them for training without buying a commercial license. It seeks at least €17.8 million.
July 15, 2026 · 404 Media report
Code comments clearly list YouTube Music and other services with hour counts. 404 Media and follow-up outlets point out this lines up with allegations that Suno scraped YouTube directly.
Why consider two layers

The first layer is copyright: is copying songs for training fair use under the law? The second is technical protection: did Suno bypass YouTube's download protections? Even if the first is unresolved, the second can stand alone. This leak primarily serves to solidify the "massive downloading from YouTube" claim.

As we have stated in public filings and disclosures, Suno’s AI models have been trained on publicly available music files and related metadata accessible on third-party websites on the open Internet. Suno spokesperson responding to 404 Media (via Music Business Worldwide)
5The Users

What the company said, and what users knew

One issue is potential copyright infringement in training data. Another is whether user data was compromised and whether users were notified. Both stem from the same intrusion.

What Suno says

Detected a limited incident in November 2025 and contained it quickly. Mainly involved superseded source code. No sensitive personal information was compromised. Suno doesn't have access to full card numbers on Stripe. Therefore, it believes it isn't legally required to notify every user. On training data, it says it's already made public disclosures per California law.

What reports and users say

Hacker ellie.191 says they could see emails, phone numbers, and Stripe-related info for several hundred thousand users. TechCrunch reported the material included some card numbers. 404 Media's Jason Koebler wrote on Bluesky that users were never notified. Some customers confirmed to 404 Media they hadn't received any breach notification and found out from the news.

Currently, two facts seem solid. First, Suno judged this a "limited incident" and didn't think individual notification was required. Second, at least some users first heard about it from media in July 2026. Whether full card numbers were exposed, and what specific fields were visible in Stripe, will depend on future statements, regulatory findings, and litigation results.

Suno's California AB 2013 disclosure page still states the training set comprises tens of millions of public music files and related text descriptions, collected since spring 2023, possibly including public domain and copyrighted works, and that it respects paywalls and passwords. What this leak added are specifics missing from previous vague statements: the specific site names, hours per library, use of Bright Data proxies, and the specific search for a cappella tracks.

The hacked data is a rare look at exactly how AI models and tools are built. Jason Koebler / 404 Media, 2026-07-15

Primary source: 404 Media original report (Jason Koebler, 2026-07-15). The latter half of 404's piece is paywalled; the free portion covers dataset hours. Details on the worm, Bright Data, a cappella searches, podcast plans, company statements, users not being notified, and the Jamendo suit were cross-referenced with Music Business Worldwide, Variety, TechCrunch, TechTimes, and the author's Bluesky. Suno's own disclosure page: California AB 2013. The nine library hour counts are summed directly; the ~1 million hours of podcasts refers to the scale of the download plan.