← Home

Frequently Asked Questions

Answers to what voice actors and buyers ask most. Still need help? Email admin@worder.com.

Getting started

How do I generate speech, and how many samples do I need?

Buyers:Open a voice from an actor’s storefront or the marketplace, click Generate Speech, type your script (at least 30 characters), and generate. You’re charged per second of audio produced, and you can hear a short free preview before you buy.

Voice actors:To create your voice, upload reference recordings of your own voice — the AI clones from these, so their quality matters far more than their quantity. A single strong sample is enough to publish a voice, but adding one sample per emotion (each clearly labeled) unlocks direction tags and a much more expressive range. Create a separate voice for each language and accent you perform. See “How should I record my voice samples for the best results?” below for detailed guidance.

How should I record my voice samples for the best results?

Your samples are the reference the AI clones from, so their quality directly shapes everything generated from your voice. A little care here makes a big difference.

One voice per language and accent

  • A voice carries one language and one accent. If you perform more than one — Spanish (Colombia) and Spanish (Mexico), or a New York and a Southern American English — create a separate voice for each. Buyers search by accent, so each one gets found on its own, and the AI clones each from samples in that accent only.
  • Samples belong to a voice and carry its language and accent. What a sample adds is emotion — the direction tags buyers use mid-script — so mixing accents inside one voice would leave a tag like [happy] with two possible answers.
  • One strong sample is enough to publish a voice. To give buyers emotional range, add one sample per emotion (see below).
  • Aim for about 10–15 seconds of clean, continuous speech per sample. Longer isn’t better — only the opening seconds are used, so extra length adds nothing. Samples over 2 minutes are rejected at upload.

One emotion per sample

  • Keep each sample to a single, consistent emotion and style from start to finish. Don’t combine, say, an upbeat line and a somber line in one recording — the model captures the overall delivery of the sample, so mixing styles muddies the result.
  • Hold your tone, pace, energy, and volume steady within each sample.
  • Label each sample with its emotion (Happy, Calm, Confident, Excited…). Those labels become the direction tags buyers use to switch emotion mid-script.

Audio quality

  • Record in a quiet space — no background noise, hum, echo, or reverb.
  • No background music, sound effects, or other voices.
  • Use a good microphone if you can (rather than a phone or laptop mic), and avoid heavily compressed or low-bitrate audio.
  • Keep a consistent, natural volume with no clipping or distortion.
  • Export in mono, or make sure both stereo channels carry audio. A file recorded from a single mic into one side of a stereo track plays out of one ear and gives the AI half a reference — those uploads are rejected.
  • Upload the highest quality you have — an uncompressed WAV is ideal.

Clean it up before uploading

  • Trim silence and dead air at the start and end.
  • Remove heavy breaths, mouth clicks, and lip smacks — but don’t strip every breath. Testers consistently name the total absence of breathing as the thing that gives an AI voice away, so light natural breaths are an asset, not a flaw.
  • Cut stumbles and false starts.

No filler words or sounds

An improvised, conversational read often sounds more natural and alive than a scripted one — that instinct is a good one, and we don’t want to flatten it. But improvising is also where “um”, “uh”, “so…” and “you know” come from, and in a reference sample those carry a cost that doesn’t exist in ordinary voice work: a filler sound in your sample does not stay in your sample.

  • The AI has no way to know a filler sound isn’t a word you meant to say. It treats everything in the sample as part of how you speak, so it can reproduce that sound in generated audio — at the start of a read, and again after every pause.
  • One actor’s sample had a soft “um” in the opening seconds. Every generation from that voice began with one.
  • The same goes for throat-clears, audible swallows, and anything you’d edit out of a finished take.

Two ways to keep the naturalness without the cost: write the words down and performthem rather than read them flat, or improvise the way you normally would and edit the fillers out before you upload. Either works — what matters is that the file you upload has none left in it.

What to say

  • Use natural, well-articulated speech in full sentences.
  • Pick content that represents the style you’re offering; avoid tongue-twisters or number-heavy text.
  • Make sure the language and accent you speak exactly match the voice this sample belongs to.

How do I use emotion (direction) tags?

If your voice has multiple labeled samples, you can switch emotion mid-script by inserting the tag before a passage:

[happy] Welcome back! [calm] Let’s take a deep breath.

Other script controls you can mix in:

  • [pause 2] — inserts an exact 2-second pause. Any length from [pause 0.5] to [pause 30] works (decimals are fine); values above 30 seconds are ignored. [pausa 3] works too.
  • [emphasize] — boosts the volume of the next phrase (great for taglines)
  • {Brand|phonetic} — controls pronunciation, e.g. {Nike|Naiki}
  • Punctuation shapes timing more subtly: commas and ellipses (…) add short natural pauses. A line break does more than pause — it changes how the model paces and inflects the whole passage, so moving a paragraph onto its own line can noticeably shift the delivery. For an exact, predictable pause, use [pause N].

How do I write a script that gets the best result?

Most of the difference between a good generation and a great one comes from the script itself. These come straight from what our testers found:

  • Use [pause N] for timing you can count on. Punctuation nudges the rhythm, but only a pause tag gives you an exact, repeatable gap.
  • Line breaks are a real creative lever. Moving a paragraph onto its own line doesn’t just add a gap — it changes how the voice paces and inflects that passage. If a read feels flat, try breaking it differently before changing the words.
  • Give every tagged section a full sentence. Each section is generated separately and needs at least 30 characters; short fragments sound clipped and are rejected.
  • Change one emotion at a time. Each direction tag is generated from a different recording, so switching emotion every line makes the delivery feel restless. Group your script into a few emotional blocks instead.
  • Keep long scripts in parts. Every tag adds generation time; very long, heavily tagged scripts can run out of time. Two shorter generations are faster and usually sound better than one sprawling one.
  • Write numbers, dates and acronyms the way they should be said (“twenty twenty-six”, “A-P-I”), and use {Brand|phonetic} for names the voice gets wrong.
  • Re-generate after small edits. Because the model reads the whole script for context, a small wording or spacing change can meaningfully shift the performance — it’s often worth a second attempt.

Are there language and accent options?

Worder currently supports 10 languages — English, Spanish, French, German, Italian, Portuguese, Japanese, Korean, Chinese, and Russian — each with regional and accent variants (for example, English or Spanish from different countries). If the language or accent you need isn’t listed yet, email admin@worder.com and we’ll keep you posted as our language support grows.

Account & access

I can't register or generate — what should I check?

  • Invite code:During the beta, signing up requires a valid invite code. Enter it exactly as given (codes are case-sensitive) and make sure it hasn’t already been used or expired. If yours doesn’t work, email admin@worder.com for a fresh one.
  • Verification codes not arriving: We email a code during sign-up and, if you enable it, at two-factor login. If it doesn’t arrive within a minute, check your spam or promotions folder and confirm you entered the right email address — then request a new code.
  • “0 matches” on your verification video:This means the words we transcribed from your recording didn’t match the phrase on screen — usually caused by background noise, speaking too quickly, or not reading the phrase exactly. Record in a quiet room, read the phrase exactly as displayed, and try again.
  • Still stuck? Email admin@worder.com with your email address and what you were trying to do, and we’ll help you directly.

Payments

How do payments work — do I need to send invoices?

No invoicing needed — everything is automated. Buyers pay per generation, and as a voice actor you earn 90% of every sale. Your balance is paid out automatically through Stripe on the 1st of each month, once it has reached $10, and you can track your earnings and payouts from your dashboard.

How do I set up Stripe so I can get paid?

Worder pays voice actors through Stripe Connect. You’ll need to do this once, before your first payout — we can’t send money to an account that hasn’t been set up.

  1. Go to your Earnings page and click Connect Stripe to get paid.
  2. Choose the country where your bank account is held. Take care here — Stripe fixes the country when the account is created and it cannot be changed afterwards. It also decides which bank fields you’ll be asked for: a routing number in the US, an IBAN across much of Europe, a sort code in the UK.
  3. Stripe will ask for your legal name, date of birth, address, and bank details. Businesses may also be asked for a tax ID.
  4. Depending on your country and earnings, Stripe may ask for a photo ID or proof of address. This is identity verification required by financial regulations — Worder never sees these documents.
  5. When you finish, you’re returned to Worder. Your Earnings page will show that payouts are enabled, usually within a few minutes.

You don’t have to finish in one sitting. If you stop halfway, the same button on your Earnings page picks up where you left off.

When do I actually receive the money?

  • Your share of each sale (90% or more) is added to your pending balance as soon as a buyer generates audio.
  • Payouts run once a month, on the 1st, and only for balances of $10 or more. Below that, your balance simply carries over to the following month.
  • Once a payout is sent, Stripe transfers it to your bank. How long that takes depends on your country and bank — typically two to seven business days for a first payout, and faster after that.
  • You can see every payout, including ones still in transit, in your Stripe Express dashboard — reachable from your Earnings page.

What should I charge for my voice?

You set your own rate, per second of audio generated, and you keep 90% of every sale. Nobody at Worder sets your price. But most voice actors are pricing for AI generation for the first time, so here’s what the market looks like.

What AI voice generation costs today

Generic AI text-to-speech — no human involved, no consent given, nobody paid — costs buyers roughly $0.10 to $0.50 per finished minute. That’s the reference point every buyer has in their head when they arrive.

Worder is more expensive than that, and it should be. Your voice belongs to a real person who agreed to the work and gets paid for it, and that’s worth paying for. But the gap has to stay reasonable: a buyer comparing you against a synthetic voice will pay a premium for a real one, not several times over.

A working range: 3¢ to 10¢ per second— about $1.80 to $6.00 per finished minute. At 5¢ a second, a minute of finished audio costs a buyer $3.00, and you earn $2.70 of it.

One thing that’s easy to miss

Buyers pay for every take. A script is rarely right the first time, and fifteen attempts at a thirty-second read is seven and a half minutes of billed audio, not thirty seconds. Your rate is charged against everything they generate, not just the version they keep — so a high rate costs you bookings in ways you never see, because the buyer quietly picks someone else instead.

The edges

  • Below about 2¢ a second you’re matching generic AI on price — a race that has no winner, and it gives away the one thing those tools can’t offer.
  • Above about 15¢ a second, buyers will want a reason they can see on your card: a rare language or accent, a distinctive character, a name they already know.

Market figures checked August 2026; AI voice pricing moves quickly.

I'm having trouble connecting Stripe, especially from outside the US.

Payouts run through Stripe Connect, which supports voice actors in many countries — but the exact bank details Stripe asks for depend on where your account is held (a routing number in the US, an IBAN across much of Europe, a sort code in the UK, and so on). A few tips:

  • At the start of onboarding, select the country where your bank account is actually held — this determines which fields you’ll see.
  • Enter your details in your country’s format. If a field like “routing number” doesn’t apply to you, choosing the correct country makes Stripe show the right equivalent instead.
  • You can return to onboarding anytime from your Earnings page if you didn’t finish.
  • A few countries aren’t yet supported by Stripe for payouts. You can check the current list at stripe.com/global. If you can’t complete onboarding, email admin@worder.com with your country and we’ll advise on options.

If you picked the wrong country

Stripe fixes an account’s country when it’s created, so it can’t be edited afterwards — the account has to be replaced. If yours is asking for bank details from the wrong country (a US routing number when you bank in Spain, say), go to your Earnings page and use “Wrong country, or asked for the wrong bank details? Start over”, pick the right country, and connect again. This deletes the old onboarding and creates a fresh one. Any earnings already in your Worder balance are unaffected.

Stripe says my account is restricted, pending, or needs more information.

This is between you and Stripe rather than something Worder controls — we can see whether your account is ready for payouts, but not what Stripe is asking you for.

  • Pending usually means Stripe is still reviewing what you submitted. Most reviews finish within a day or two.
  • Restricted or information needed means something is missing — often an ID document, a proof of address, or a bank account that couldn’t be verified. Stripe will list exactly what it needs.
  • Open your Stripe Express dashboard from your Earnings page to see the outstanding items and upload what’s missing. You can also reach it directly at connect.stripe.com/express_login.
  • For anything Stripe-specific — a rejected document, a verification that won’t clear, a question about their requirements — Stripe’s own support is the fastest route: support.stripe.com.

Your earnings keep accruing on Worder while this is being sorted out. Nothing is lost — payouts simply wait until Stripe enables them, then send as normal.

Ownership, privacy & safety

What control do I have over my voice and identity?

You’re always in control. You can edit or remove your voices and samples at any time, and set usage restrictions — content categories your voice may not be used for — in your account settings.

If you delete your account, your voice models and samples are permanently removed and can never be used to generate new audio again. Audio that a buyer already generated and downloaded before deletion remains theirs under the license they purchased — we can’t recall files that were already delivered — but nothing new can ever be produced from your voice. To request account deletion, email admin@worder.com.

What safeguards exist against misuse?

  • No AI training on your voice: Worder does not use your samples to train, fine-tune, or improve AI models, and buyers are contractually prohibited from training on the audio they generate.
  • Automated screening:Every generation request is screened before any audio is produced. Content in the categories you’ve restricted, impersonation of real people, and content targeting or harming real individuals are automatically blocked.
  • Verified, credited voices:Every voice is tied to a verified, real person. This is core to Worder being a fair-trade marketplace — buyers know exactly whose voice they’re licensing, and actors are properly credited and paid. Voices are not anonymous.

Quality & what’s next

How can I improve the quality of my AI clone?

Use clean, expressive reference samples recorded in a quiet room, and take advantage of the script tools — direction tags ([happy], [calm]), pauses ([pause 2]), emphasis ([emphasize]), and pronunciation overrides ({Brand|phonetic}) — to shape the delivery. Natural sentence lengths and clear punctuation also help the model produce more natural-sounding speech.

When does Worder launch, and is there a tutorial?

The voice-actor beta is live, and access for buyers is opening in beta soon. A video walkthrough is on the way — we’ll share it as soon as it’s ready. In the meantime, if you’d like a hand getting set up, email admin@worder.com.

Didn’t find your answer? Email admin@worder.com and we’ll help.