Audio is where the copyright and identity questions hanging over all of generative AI have come to a head first and hardest, because a voice and a song are so personal. The image and video labs will face versions of this; the audio companies are living it now. Two threads run through everything in this part: who owns the training data, and who owns a voice.
01
The turn toward licensed data
The dominant story of 2025 and 2026 is a decisive pivot toward licensed or owned training data as the basis for commercial safety, and two camps emerged. The first is the licensed-data pure plays, Stability and ElevenLabs, who trained on rights-cleared catalogs (Stable Audio on licensed and Creative Commons data, Eleven Music under opt-in deals with Merlin and Kobalt) and market indemnificationIndemnificationA vendor's promise to cover a customer's legal costs if the product triggers a lawsuit, offered by some licensed-data model providers as protection against copyright claims on generated output. and clean provenanceProvenanceA verifiable record of how a piece of media was made and by whom, increasingly required so that synthetic content can be traced. as the product. The second is the litigate-then-license path of the incumbents: Suno and Udio were sued by the major labels, then converted those suits into licensing deals, Udio by settling with Universal, Warner, Merlin, and Kobalt into a walled gardenWalled gardenA platform where content can be created, customized, and shared internally but never exported, the position Udio took after its licensing deals., Suno by settling with Warner while fighting to keep downloads. Critically, the settlements are not universal. Sony has settled with neither, and Universal's case against Suno was unresolved as of mid-2026, so a definitive US fair-use ruling could still land in 2027 and reshape everything.
02
A voice is now legally you
The law is racing to treat a voice as a protectable part of a person's identity, with consent as the organizing principle. Tennessee's ELVIS Act, effective July 2024, was the first state law to explicitly protect a voice, including simulations of it, and notably it reaches the makers of the cloning tools, not just the end users. The federal NO FAKES ActNO FAKES ActA proposed United States federal law that would create a nationwide right against unauthorized digital replicas of a person's voice and likeness., which would create a nationwide right against unauthorized digital voice and likeness replicas, has been introduced repeatedly but not enacted, and it is genuinely contested, with critics warning its takedown-style obligations could burden lawful expression. On the labor side SAG-AFTRA has moved fastest and most concretely, striking deals that require informed consent both to create a replica and, separately, again to use it, plus compensation, time limits, and disclosure.
Fig.diagram
■The two-consents rule, one consent to create a replica and a separate consent for every use.
03
Provenance, and the through-line
The technical answer to all this is provenanceProvenanceA verifiable record of how a piece of media was made and by whom, increasingly required so that synthetic content can be traced., and it comes in two flavors that are often confused. Embedded watermarks, like Google's SynthIDSynthIDGoogle's watermarking system, which embeds an identifying signal in images, audio, video, and text it generates., Resemble's PerTH, and Meta's AudioSealAudioSealMeta's audio watermarking method, designed to embed a signal detectable down to short segments and to survive light editing., hide a signal inside the audio itself that survives some editing and supports detection. Content Credentials, the C2PAC2PA (Content Credentials)An open provenance standard, backed by Adobe, OpenAI, and Microsoft, that attaches cryptographically signed metadata proving a file's origin, though the metadata strips easily when a file is re-encoded. standard backed by Adobe, OpenAI, and Microsoft, attach cryptographically signed metadata that proves origin when present but strip easily when a file is re-encoded. Both help; neither is a guarantee. Recent research is blunt that watermarkingWatermarkingHiding a detectable signal inside generated media so its origin can later be identified. For audio and AI media it remains an arms race rather than a guarantee. and detection remain an arms race that can be gamed, so provenance should be treated as a deterrent and an audit aid, not proof.
Check yourself0 / 6
Q01
What are the two camps in the pivot toward licensed data?
Q02
Why do the music settlements not yet establish anything legally?
Q03
What made Tennessee's ELVIS Act notable?
Q04
What is the two-consents principle, and why does it surprise creators?
Q05
How do embedded watermarks differ from Content Credentials, and what is the honest verdict on both?
Q06
What is the recurring through-line of the audio reckoning?
The weekly briefing
Get the week's moves in your inbox.
A short, sourced digest of what actually moved across generative AI, every week. Free.
Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.