Skip to content
Public Art Now

Free AI Voice Generators, and Where You Cannot Use Them

Public Art Now featured card reading Free AI Voice Generators and Where They Stop, beside a microphone with sound arcs and a MARKED badge

Synthetic narration crossed the believability line some time ago. The remaining problems are not acoustic, they are legal, and they are easy to miss because nothing in the interface mentions them. The most popular free voice tier grants no commercial rights whatsoever, the best free cloning model stamps an inaudible watermark into every file it produces, and the only genuinely unencumbered option cannot clone a voice at all.

Is ElevenLabs free version usable for commercial work?

No. ElevenLabs is the best-known name here and its pricing page gives free accounts 10,000 credits a month across text to speech, speech to text, sound effects, voice design and music, plus three Studio projects. It is a generous sandbox.

The commercial licence begins at the Starter tier, $6 a month, and so does voice cloning. Nothing produced on the free plan is cleared for a monetised video, a client project, or a product. Free here means practise, not publish.

That is a defensible arrangement and a cheap upgrade. The problem is how many people never read the tier and assume a downloadable MP3 with no visible watermark is theirs to use.

Chart placing four AI voice options by what the licence permits, from personal use only to use it anywhere with no conditions
Audio quality is not what separates these. Permission is. Positions read from each service’s pricing page or model card at the time of writing.

Kokoro, small, free and completely unencumbered

Kokoro-82M is the quiet answer for anyone who needs narration rather than impersonation. Its model card puts it at 82 million parameters under Apache 2.0, and claims quality comparable to much larger models while being significantly faster and more cost-efficient.

The economics are startling. The same card puts the market rate for Kokoro served over an API at under $1 per million characters of text input, or under $0.06 per hour of audio output, and notes the model was trained for roughly $1,000 in A100 compute. Run locally it costs nothing at all.

What it does not do is clone. Kokoro ships fixed voices, which is precisely why it raises no consent question and no rights question. For audiobooks, explainer videos, accessibility narration and interface speech, that limitation is not a limitation.

A model that cannot clone anybody is the only one that never has to ask anybody’s permission. That is a feature.

Chatterbox, open cloning that marks its own work

Chatterbox from Resemble AI is the most capable free option, and the most interesting. Its model card describes production-grade open source text to speech under the MIT licence, with multilingual support, emotion exaggeration control, and zero-shot voice cloning from a supplied audio prompt.

MIT means no commercial restriction at all. But the card documents something unusual for an open release: every audio file it generates carries Resemble AI’s Perth perceptual watermarker, described as imperceptible neural watermarks that survive MP3 compression, audio editing and common manipulations while maintaining nearly 100 per cent detection accuracy.

Read that carefully, because it cuts both ways. The output is permissively licensed and permanently identifiable as synthetic. For legitimate work that is a benefit, since provenance is increasingly something platforms and clients ask about. For anyone hoping to pass synthetic speech off as a real recording, it is a wall, and deliberately so.

Quick comparison

ToolFree allowanceCommercial useVoice cloning
ElevenLabs, free10,000 credits a monthNo, Starter tier and aboveNo, Starter tier and above
Kokoro-82MUnlimited, run locallyYes, Apache 2.0No, fixed voices only
ChatterboxUnlimited, run locallyYes, MITYes, zero-shot, watermarked

Where you genuinely cannot use a cloned voice

The licence on the model is only the first of three gates, and it is the easiest one to clear. A permissive licence grants what the model’s authors can grant. It says nothing about the person whose voice you copied.

  • Someone else’s voice, without written permission. Voice is treated as a personal attribute in many jurisdictions, protected separately from copyright. An MIT licence on the software is not consent from the speaker.
  • Anything implying endorsement. A synthetic voice recommending a product reads as that person recommending it, which is where advertising and consumer-protection rules apply regardless of how the audio was made.
  • Undisclosed synthetic speech in regulated contexts. Political advertising, financial promotion and automated calling carry disclosure requirements in a growing number of places, and they attach to the use rather than the tool.
  • Platform-restricted uploads. Several services require synthetic speech to be labelled, and impersonation of a real person is prohibited outright on most. Legal and permitted are different tests.

The safe practice is unglamorous. Clone your own voice, use a fixed synthetic voice, or get written permission with the intended use spelled out. Anything else puts the risk on you and not on the model’s licence.

Getting better output from any of them

Synthetic speech fails in predictable places, and most of those failures are fixable in the text rather than the settings.

  1. Punctuate for breath, not for grammar. A comma is where the voice pauses. Adding them against the rules of written English usually improves delivery.
  2. Spell out anything unusual. Numbers, acronyms, place names and units are where models mispronounce most often. Write them phonetically in the source text.
  3. Generate in paragraphs, not pages. Long passages drift in pace and tone. Shorter chunks stitched together stay consistent.
  4. Give a clean cloning sample. Zero-shot cloning copies whatever it hears, including room echo and background noise. Thirty seconds of clean audio beats five minutes of noisy audio.

The bottom line

Kokoro is the right default for narration, because Apache 2.0 removes every question and the running cost is effectively zero. Chatterbox is the right choice when a specific voice is required and you have the right to use it, with the watermark as a fair condition. ElevenLabs remains the best-sounding option, and its free tier is a rehearsal space rather than a licence.

The question that decides all of this is not which tool sounds best. It is whose voice is coming out of it, and whether that person agreed. Settle that first and the rest is a five-minute configuration problem.

Frequently asked questions

Can I use ElevenLabs free audio on YouTube?

Not on a monetised channel. The free plan provides 10,000 credits a month but no commercial licence, which begins at the $6 Starter tier. A monetised upload is commercial use even when the narration is incidental to the video.

Which free AI voice generator allows commercial use?

Kokoro-82M under Apache 2.0 and Chatterbox under MIT both permit commercial use outright, and both run locally at no cost. Kokoro offers fixed voices only; Chatterbox adds zero-shot cloning but watermarks every file it produces.

Do AI voice tools watermark their output?

Chatterbox does. Its model card states every generated file carries Resemble AI’s Perth neural watermark, imperceptible to listeners, surviving MP3 compression and editing with nearly 100 per cent detection accuracy. It marks the audio as synthetic rather than restricting its use.

Not without permission, in most places. Voice is protected as a personal attribute separately from copyright, so a permissive licence on the model grants nothing regarding the speaker. Clone your own voice, use a fixed synthetic voice, or obtain written consent that names the intended use.

How much does AI text-to-speech cost to run?

Very little. Kokoro’s model card puts the market rate for the model served over an API at under $1 per million characters of input, or under $0.06 per hour of audio produced. Run on your own machine, the marginal cost is zero.

Where to Download Museum Art Images Legally

AI & Creative Tools 7 min

Hundreds of thousands of high-resolution artworks are free to download and reuse commercially, straight from the museums that hold the originals.