Focused server-side examples for building expressive text-to-speech workflows with Gemini 3.1 Flash TTS on PoYo.
Gemini 3.1 Flash TTS is useful for expressive narration, multi-speaker dialogue drafts, audio tags, product voiceovers, and creator-tool speech generation.
Try on PoYo | Get API Key | Docs | Pricing | Main Examples
- Text-to-speech with
gemini-3-1-flash-tts - Style instructions and voice control
- Output format selection
- Async task flow with polling and webhooks
- cURL and Node.js backend examples
cp .env.example .env
export POYO_API_KEY="your-api-key"Run the Node.js example:
cd node
npm startKeep POYO_API_KEY on the server. Do not expose it in browser code, mobile apps, screenshots, or public logs.
- Keep
POYO_API_KEYon the server - Submit a generation task
- Store
data.task_id - Poll status while testing
- Use
callback_urlwebhooks in production
This repo uses gemini-3-1-flash-tts.
| Path | What it covers |
|---|---|
curl/generate.md |
Copy-paste API request. |
node/ |
Native Node.js backend example. |
docs/prompt-examples.md |
Practical prompts for product workflows and creative tests. |
docs/production-notes.md |
Security and reliability notes before launch. |
webhooks/express-webhook/ |
Minimal Express receiver for PoYo callbacks. |
make checkOn Windows:
./scripts/check.ps1MIT