मुख्य सामग्री पर जाएं

Text-to-Speech Studio

Creative skill, available on Zeplik

Text-to-Speech Studio is a ready-to-run creative skill on Zeplik. Spoken audio from text, which the product does synthesize (paid, inline player). Ask in plain language and Zeplik applies the skill's method for you inside the conversation, on whichever AI model you prefer.

The Text-to-Speech Studio skill loads automatically when your request matches it, or you can invoke it directly by typing /text-to-speech in any chat. It works with attachments, connectors, and any model that supports the task, so you get the same expert method every time without setting anything up.

What the Text-to-Speech Studio skill can do

Try these prompts on Zeplik

Pick a prompt to open it in the Zeplik app. If you are not signed in yet, your prompt is waiting for you the moment you do.

How the Text-to-Speech Studio skill works

/text-to-speech

Capability truth (read first): this product DOES turn text into real spoken audio. It is a paid generation, it lands as an inline audio player the user can play and download, and the engine name and the price are stated in your system prompt — read them there, never quote either from this file. So the answer to "can you make a voiceover / narrate this / give me an audio version" is yes. Never deny it, and never send the user to another service for it.

Why you are reading this, and who actually presses the button

Synthesis is dispatched by the router, ahead of you — you will not find a speech tool in your tool list and there is nothing here for you to invoke. The router reads the user's own words; when it recognises a speech request it owns the whole turn and you are never called. So the mere fact that this skill is in front of you means the router read this turn as something other than a speech command. That is a fact about phrasing. It is not evidence that the capability is missing, and you must not report it as such.

That leaves you with the two jobs that are genuinely yours, and one handover:

  1. Say yes. Confirm plainly that spoken audio is something the product makes.
  2. Do the craft. Prepare the voiceable script and name the voice (below).
  3. Hand over. The phrasing that dispatches names speech output explicitly — "read that aloud", "read this aloud: <text>", "say it out loud", "turn this into speech", "text to speech", or asking for the audio file by name ("give me an mp3 of that"). Tell the user, in one short line, to say it that way — or to pin Speech in the composer, which declares the intent outright and voices the whole message. One line. It is a handover, not an apology.

The prior-reply path voices your ENTIRE last message. A bare "read that aloud" synthesizes the whole of your previous reply verbatim — headings, preamble, the voice-direction block, all of it. So when the user is about to have your text voiced, the final reply must be the script and nothing else. Put any commentary, direction notes or options in an earlier turn or above a clear break the user can trim; do not leave them sitting in the message that gets read.

What the engine actually honors

Steer only these. Everything else is craft that lives in the script itself:

  • Voice — a roster of named voices; the user selects one by naming it against the word "voice" ("in Leo's voice", "the Carina voice", "voice: aurora") or after a bare selection verb with no article ("using Ara", "as Luna", "voiced by Rex"). The article matters: "read it in Atlas" selects a voice, "read it in the atlas" reads a book. A tone or gender cue ("a deep voice", "a woman's voice") also selects one. A voice merely mentioned does not.
  • Language — an explicit ask ("read this in Spanish", "in Japanese"). Absent one, the engine detects the text's own language, which is the right default for reading arbitrary content back.
  • Engine — the user may name an exposed speech engine, or pick it with the composer's Model chip.

A speed / pacing / emotion parameter is NOT honored — this was probed against the live engine, and a "pacing: slow" line reaches nothing. The only pacing control that actually changes the audio is the script's own punctuation and line breaks, which is exactly why script prep below is the load-bearing craft here and not a formality. Say so honestly if a user asks for a slower read: the way to get one is to rewrite the sentences shorter, not to add a direction line.

Length. Long text is truncated at the engine's cap rather than refused, and the cap is lower on some engines than others — so split a long piece into sections at natural boundaries and voice them one at a time, keeping the same voice across sections. Do not hand over a wall of text and hope.

Not the same feature: read-aloud and Voice mode

The platform also has a read-aloud control and Voice mode that speak a reply through the UI. That is playback — no saved file, nothing charged, nothing the user can download. When someone wants a file (a voiceover to drop into an edit, an mp3 to publish), that is the generation path above, not this one. Offer the right one; do not substitute playback for a deliverable.

Workflow

  1. Collect inputs up front: the exact text (verbatim), the voice or delivery the user wants, and where the audio will be used (demo voiceover, narration, phone prompt, accessibility read).
  2. Prepare the script for the ear (only with the user's consent if it changes their words):
    • Break long sentences; spoken clauses should fit in one breath. This is the real pacing lever.
    • Expand or hint pronunciations: acronyms as letters ("A-I"), tricky names with a simple phonetic guide in the text.
    • Numbers, dates, and URLs: write them the way they should be spoken ("twenty twenty-six", "zeplik dot ai").
    • Use punctuation and short line breaks to create natural pauses.
    • Cut visual-only phrasing ("as shown below", "see the table above") and anything that is markup rather than speech.
  3. Name the voice if the user expressed any preference, in the selecting form the parser reads ("in Leo's voice", "using Ara") so their next turn carries it.
  4. Deliver the script as the reply — alone, if it is about to be voiced — and add the one-line handover.
  5. Guide validation: what to listen for on playback — intelligibility, pacing, pronunciation, names.
  6. Iterate with a single targeted change (the script's phrasing, or the voice); repeat invariants to reduce drift.
  7. For long texts, split at section boundaries and keep the voice identical across sections.

Voice direction template

Useful for two things: shaping the script (pauses, emphasis and pronunciation all become punctuation and word choice), and writing a spec for an external TTS engine that does take direction. Our engine takes voice and language only — so never present these lines to the user as knobs on the audio they are about to get. Include only relevant lines, 4 to 8 total:

Voice Affect: <overall character and texture of the voice>
Tone: <attitude, formality, warmth>
Pacing: <slow, steady, brisk>
Emotion: <key emotions to convey>
Pronunciation: <words to enunciate or emphasize>
Pauses: <where to add intentional pauses>
Emphasis: <key words or phrases to stress>
Delivery: <cadence or rhythm notes>

Example

Input text: "Welcome to the demo. Today we'll show how it works."
Direction:
Voice Affect: Warm and composed.
Tone: Friendly and confident.
Pacing: Steady and moderate.
Emphasis: Stress "demo" and "show".

Direction best practices (short list)

  • Order: affect -> tone -> pacing -> emotion -> pronunciation/pauses -> emphasis.
  • Prefer concrete guidance over adjectives alone; avoid conflicting lines ("fast and slow").
  • Do not rewrite the input text inside the direction; the direction guides delivery only.
  • If the user needs another language, write the input text in that language and ask for that language explicitly.
  • Disclose to end listeners that the voice is AI-generated when the audio will be published.

Guardrails

  • Never fake a result. No invented audio blob, no {"audio": …} JSON, no data URL, no placeholder link, and never say a reading "is attached" — the audio arrives as a real player from the generation path or it does not exist yet.
  • Never quote the engine name or the price from this file; the system prompt is the only current source.
  • No real people's voices and no voice cloning: the roster is the roster. There is no audio-to-audio voice conversion — you cannot restyle a clip someone uploaded, only synthesize fresh speech from text.

Guidance by use case

  • Narration / explainer: references/narration.md
  • Product demo / voiceover: references/voiceover.md
  • IVR / phone prompts: references/ivr.md
  • Accessibility reads: references/accessibility.md

More references

  • references/voice-directions.md -- direction template + styled examples (calm support, dramatic narrator, energetic, serene, robotic, announcer).
  • references/prompting.md -- direction-writing principles and iteration patterns.
  • references/sample-prompts.md -- copy/paste direction blocks.

Usage

/text-to-speech $ARGUMENTS

How to use the Text-to-Speech Studio skill

  1. Sign in to Zeplik

    Create a free Zeplik account or sign in. New accounts start with free credits, so you can try the Text-to-Speech Studio skill right away.

  2. Describe your creative task

    Ask in plain language, or type /text-to-speech to invoke the skill directly. Zeplik recognizes the Text-to-Speech Studio skill and applies its method.

  3. Review and refine the result

    Zeplik returns a clear, structured answer. Ask follow-ups in the same chat to refine it or take the next step.

Source and credit

Author
davila7
License
MIT

Adapted from the open-source davila7/claude-code-templates project and tuned to run natively on Zeplik. View source on GitHub.

Frequently asked questions

What is the Text-to-Speech Studio skill?
Text-to-Speech Studio is a ready-to-run creative skill on Zeplik. Spoken audio from text, which the product does synthesize (paid, inline player). Ask in plain language and Zeplik applies the skill's method for you inside the conversation, on whichever AI model you prefer.
How do I use Text-to-Speech Studio on Zeplik?
Sign in to Zeplik and ask in plain language, or type /text-to-speech in any chat to invoke it directly. The skill applies its method and returns a result you can refine in the same conversation.
Which AI model does the Text-to-Speech Studio skill use?
Any model you choose. Zeplik works across every model in one chat, so the Text-to-Speech Studio skill runs on your preferred model for the task.
Where does the Text-to-Speech Studio skill come from?
The Text-to-Speech Studio skill is adapted from the open-source davila7/claude-code-templates project (MIT) and tuned to run natively on Zeplik. The original source is linked on this page.
How much does the Text-to-Speech Studio skill cost?
Using the skill is free to start. You only spend Zeplik credits when the assistant runs, and new accounts begin with free credits.

Related creative skills

More on Zeplik

Try Text-to-Speech Studio on Zeplik

Every model, one chat. Bring the Text-to-Speech Studio skill into your next conversation and let the assistant do the work.

Browse all skills