Emotion and speaking-style control, with asynchronous jobs
Terms of provision
token header (64-character access token) as the Standard TTS API. For polling, send the job_token returned at job creation in the X-Job-Token header (Overview and authentication)The Advanced TTS API is a more natural-sounding speech synthesis endpoint that supports emotional expression and speaking-style control. It uses asynchronous jobs: you POST a request to get a job ID, then poll for the result.
Current availability: The Advanced TTS API is offered as a Beta. The public endpoints are POST /api/advanced-tts/ (job creation) and GET /api/advanced-tts/jobs/{job_id}/ (polling). Authenticate with the same token header as the Standard TTS API, or with a logged-in session cookie plus a CSRF token.
The samples below submit a job with the token header. To use the same session authentication as the Web UI, send the logged-in cookie and X-CSRFToken.
| Header | Value |
|---|---|
Content-Type | application/json |
token | Access token (64 characters). No CSRF token is needed with token authentication |
X-Job-Token | Job access token. When polling with GET, send the job_token from the POST response in this header |
Cookie | sessionid=...; csrftoken=... (session authentication) |
X-CSRFToken | The same value as the csrftoken cookie (session authentication) |
Referer | https://ondoku3.com/en/advanced-tts-beta/ (session authentication) |
| Field | Type | Required | Description |
|---|---|---|---|
text | string | Required | Text to read aloud (at least 10 characters; the API accepts up to 400 characters on the free plan and 3,000 characters on paid plans) |
voice | string | Optional | Voice name (choose from the list below). Default: Ruby |
tone | string | Optional | Speaking style (e.g. warm and friendly, storytelling). Up to 100 characters |
speed | number | Optional | 0.5 to 1.5 (default 1.0) |
pitch | number | Optional | -1.0 to 1.0 (default 0.0) |
model | string | Optional | flash (fast) / pro (high quality) |
seed | integer | Optional | Random seed that keeps the voice consistent (-2147483648 to 2147483647) |
Polling with the ?token={job_token} query parameter is not supported. Always send the job status request with the X-Job-Token header.
# 1) Submit a TTS job
JOB=$(curl -s \
-X POST https://ondoku3.com/api/advanced-tts/ \
-H "Content-Type: application/json" \
-H "token: $ONDOKU_TOKEN" \
-d '{"text":"Hello, this is Ondoku.","voice":"Ruby","tone":"warm and friendly"}')
JOB_ID=$(echo "$JOB" | jq -r .job_id)
JOB_TOKEN=$(echo "$JOB" | jq -r .job_token)
sleep "$(echo "$JOB" | jq -r '.min_poll_after_ms / 1000')"
# 2) Poll until the job finishes
while :; do
R=$(curl -s \
-H "X-Job-Token: ${JOB_TOKEN}" \
"https://ondoku3.com/api/advanced-tts/jobs/${JOB_ID}/")
STATUS=$(echo "$R" | jq -r .status)
echo "status: $STATUS"
[ "$STATUS" = "succeeded" ] && break
[ "$STATUS" = "failed" ] && { echo "$R"; exit 1; }
sleep "$(echo "$R" | jq -r '(.poll_after_ms // 3000) / 1000')"
done
# 3) Download the MP3
URL=$(echo "$R" | jq -r .url)
curl -L "$URL" -o output.mp3
import os
import time
import requests
TOKEN = os.environ["ONDOKU_TOKEN"]
# 1) Submit a job
res = requests.post(
"https://ondoku3.com/api/advanced-tts/",
headers={
"Content-Type": "application/json",
"token": TOKEN,
},
json={
"text": "Hello, this is Ondoku.",
"voice": "Ruby",
"tone": "warm and friendly",
},
timeout=30,
)
res.raise_for_status()
job = res.json()
job_id, job_token = job["job_id"], job["job_token"]
time.sleep(job["min_poll_after_ms"] / 1000)
# 2) Poll
while True:
r = requests.get(
f"https://ondoku3.com/api/advanced-tts/jobs/{job_id}/",
headers={"X-Job-Token": job_token},
timeout=10,
).json()
print("status:", r["status"])
if r["status"] == "succeeded":
break
if r["status"] == "failed":
raise RuntimeError(r)
time.sleep(r.get("poll_after_ms", 3000) / 1000)
# 3) Download the MP3
mp3 = requests.get(r["url"], timeout=60).content
with open("output.mp3", "wb") as f:
f.write(mp3)
print(f"saved: output.mp3 ({len(mp3)} bytes)")
import { writeFile } from "node:fs/promises";
const TOKEN = process.env.ONDOKU_TOKEN!;
// 1) Submit a job
const submit = await fetch("https://ondoku3.com/api/advanced-tts/", {
method: "POST",
headers: {
"Content-Type": "application/json",
token: TOKEN,
},
body: JSON.stringify({
text: "Hello, this is Ondoku.",
voice: "Ruby",
tone: "warm and friendly",
}),
});
type AdvancedTTSJobResponse = {
job_id: string;
job_token: string;
status: "pending" | "running" | "succeeded" | "failed";
min_poll_after_ms: number;
poll_url: string;
};
type AdvancedTTSPollResponse = {
job_id: string;
status: "pending" | "running" | "succeeded" | "failed";
poll_after_ms?: number;
url?: string;
pk?: number;
seed?: number;
code?: string;
error?: string;
};
const job = (await submit.json()) as AdvancedTTSJobResponse;
// 2) Poll
let result: AdvancedTTSPollResponse;
await new Promise((resolve) => setTimeout(resolve, job.min_poll_after_ms));
while (true) {
const r = await fetch(`https://ondoku3.com/api/advanced-tts/jobs/${job.job_id}/`, {
headers: { "X-Job-Token": job.job_token },
}).then((res) => res.json() as Promise);
console.log("status:", r.status);
if (r.status === "succeeded") { result = r; break; }
if (r.status === "failed") throw new Error(JSON.stringify(r));
await new Promise((resolve) => setTimeout(resolve, r.poll_after_ms ?? 3000));
}
// 3) Download the MP3
if (!result.url) throw new Error("audio URL is missing");
const mp3 = Buffer.from(await (await fetch(result.url)).arrayBuffer());
await writeFile("output.mp3", mp3);
console.log(`saved: output.mp3 (${mp3.byteLength} bytes)`);
POST response (202 Accepted):
{
"job_id": "YOUR_JOB_ID",
"job_token": "YOUR_JOB_TOKEN",
"status": "pending",
"min_poll_after_ms": 4550,
"poll_url": "/api/advanced-tts/jobs/YOUR_JOB_ID/"
}
| Field | Type | Description |
|---|---|---|
job_id | string | Job ID (UUID) |
job_token | string | Job access token required to poll this job (different from the API access token in the token header) |
status | string | Status right after submission. Usually pending |
min_poll_after_ms | number | Recommended wait in milliseconds before the first poll |
poll_url | string | Relative URL to poll. When you access it, send the job_token from the POST response in the X-Job-Token header |
GET polling (in progress):
{
"job_id": "YOUR_JOB_ID",
"status": "pending",
"poll_after_ms": 1000
}
GET polling (succeeded):
{
"job_id": "YOUR_JOB_ID",
"status": "succeeded",
"url": "https://storage.googleapis.com/ondoku3/voice/...",
"pk": 12345,
"seed": 123456789
}
GET polling (failed):
{
"job_id": "YOUR_JOB_ID",
"status": "failed",
"code": "internal_error",
"error": "内部サーバーエラーが発生しました"
}
| Field | Type | Description |
|---|---|---|
job_id | string | Job ID (UUID) |
status | string | pending / running / succeeded / failed |
poll_after_ms | number | Returned while in progress: recommended wait in milliseconds before the next poll |
url | string | Returned only on success: URL of the audio file |
pk | number | Returned only on success: ID of the generation history entry (Sentence) |
seed | number | Returned only when a seed is stored: the random seed used for generation |
code | string | Returned only on failure: error code (see Errors and notice responses) |
error | string | Returned only on failure: error message (currently in Japanese) |
The 25 voice names you can specify in the voice parameter of the Advanced TTS API. You can listen to each voice in the voice sample article.
| Voice name |
|---|
Anna |
Chloe |
Ellis |
Emma |
Flora |
Iris |
Lena |
Luna |
Misa |
Ruby |
Sophie |
Tina |
| Voice name |
|---|
Ash |
Chris |
Eden |
Gray |
Hope |
Hugo |
Kai |
Leo |
Noah |
Reid |
Roy |
Sam |
Yann |
Errors are returned as JSON with code (error code) and error (message), except for 403 CSRF errors and 405. Branch on code, not only on the HTTP status; the error message is currently in Japanese. Failures after a job is submitted are returned as status=failed in the polling result. The Web UI plays a notice audio depending on code, but API responses do not include a notice audio URL.
POST /api/advanced-tts/| HTTP | code | Cause |
|---|---|---|
| 400 | invalid_token | The token header is not 64 characters, or no active user matches it |
| 400 | invalid_json | The body cannot be parsed as JSON |
| 400 | validation_error | text is shorter than 10 characters, consists only of emoji, or repeats the same character 20 or more times in a row |
| 400 | invalid_speed / invalid_pitch | Not a number, or out of range |
| 400 | invalid_temperature | A temperature field (not covered by this document; you do not need to send it) was sent, and its value is not a number or is not one of 0 / 0.5 / 0.7 / 1.0 |
| 400 | invalid_seed | seed is not an integer, or out of range |
| 400 | invalid_model | model is not flash / pro |
| 400 | client_ip_unresolved | The client IP address could not be determined for a request without a token and without login |
| 400 | text_too_long | Exceeds the accepted length (400 characters on the free plan / 3,000 characters on paid plans) |
| 400 | invalid_voice | Unknown voice name |
| 400 | content_policy_violation | The text cannot be read aloud under the usage policy |
| 400 | invalid_request | Other input errors (for example, the text exceeds the limit after user dictionary replacements) |
| 403 | (HTML) | Session authentication without a token header, and the CSRF token does not match or you are not logged in |
| 405 | (no body) | A method other than POST |
| 429 | quota_exceeded | Not enough characters available. available (remaining characters) is also returned |
| 429 | rate_limited | Rate limit exceeded (Character limits and rate limits). Wait retry_after seconds (the Retry-After header has the same value) and retry |
| 500 | service_not_configured | Server-side misconfiguration |
| 500 | internal_error | Server-side error (retry after a while) |
| 503 | unsupported_location | Processing from a region where speech synthesis cannot be provided. For compatibility, error_code (UNSUPPORTED_LOCATION) is also returned |
| 503 | tts_retry_exhausted | Temporarily unable to accept the request (retry after a while) |
GET /api/advanced-tts/jobs/{job_id}/| HTTP | code | Cause |
|---|---|---|
| 200 | status=failed | The job failed. Check status / code in the JSON, not the HTTP status (codes are in the table below) |
| 403 | forbidden | X-Job-Token does not match the job (with a logged-in session: the job is not yours) |
| 404 | not_found | The job does not exist |
| 405 | (no body) | A method other than GET |
status moves from pending (waiting) → running (processing) → succeeded (done) or failed (failed). When it is failed, code is one of the following. Speech generation may also return codes not listed here, so treat unknown codes as failures too.
| code | Cause |
|---|---|
content_policy_violation | The text cannot be read aloud under the usage policy. Change the text and resubmit |
unsupported_location | Processing from a region where speech synthesis cannot be provided |
tts_retry_exhausted | Temporarily unable to generate. Resubmit after a while |
invalid_request | Input error during generation |
job_timeout | Generation took too long and was stopped. Resubmit |
enqueue_failed | The job could not be queued |
user_unavailable | The account was withdrawn or deleted during generation |
internal_error | Server-side error |
| Plan | Accepted by the API |
|---|---|
| Free | 400 characters |
| Paid | 3,000 characters |
Note: Even on paid plans, the Web UI input field shows up to 1,000 characters by default (the Beta setting for longer text extends it to 3,000 characters). Regardless of that setting, the API accepts up to 3,000 characters on paid plans.
Currently enforced: Job creation POST /api/advanced-tts/ is limited to 10 requests per 60 seconds per client IP address (shared across all tokens used from the same IP). When you exceed it, the API returns HTTP 429 with rate_limited, and retry_after and the Retry-After header contain the number of seconds to wait.
In addition, dedicated rate limits for Advanced TTS are currently being observed in shadow mode. These dedicated limits do not block requests yet and do not return 429 when exceeded. Once enforced, exceeding the values below will return rate_limited JSON and a Retry-After header.
| Target | Counted per | Planned limit | Notes |
|---|---|---|---|
Job creation POST /api/advanced-tts/ | User ID (IP address when not logged in) | 30 requests / 60 seconds 180 requests / 1 hour | Not split by user status such as paid or free |
Status check GET /api/advanced-tts/jobs/{job_id}/ | Logged-in user, or job access token + job_id | 120 requests / 60 seconds | Prevents excessive polling of the same job |
| Total status checks | Logged-in user, or job access token | 600 requests / 60 seconds | Prevents polling many jobs at once. When a job created with API token authentication or a logged-in session is checked with the X-Job-Token header, it is counted per job access token, not per owner user |
| Checks of nonexistent or unauthorized jobs | IP address | 30 requests / 60 seconds | Prevents brute-forcing job_id values |
Insufficient characters are returned as quota_exceeded JSON. Hitting the current general rate limit returns rate_limited JSON. The Web UI plays a notice audio, but API clients should branch on code in the JSON, not only on the HTTP status.
Example 429 response (the error message is returned in Japanese):
{
"code": "rate_limited",
"error": "レート制限を超えました。しばらく経ってから再試行してください。",
"retry_after": 55
}
For plan details and prices, see the pricing page; for terms of use, see the Terms of Service.
| HTTP / situation | What to do |
|---|---|
| 429 (Rate limit) | Wait retry_after seconds, or at least 5 seconds, then retry |
| 5xx | Exponential backoff (2 seconds → 4 seconds → 8 seconds, up to 3 times) |
| 4xx (such as 400 and 403) | Do not retry. Check the CSRF token and the voice name |
| Job status=failed | Check code. A failed generation may fail again when retried with the same seed, so retry with a different seed |
| Job stays pending for a long time | If it stays pending for 5 minutes or more, it may be stuck on the server. Submit it again as a new job |
min_poll_after_ms in the POST response as the minimum waitpoll_after_ms, treat it as the minimum wait before the next pollpoll_after_ms, poll every 3 to 5 seconds. Polling too often may hit the rate limiturl in the response is a signed, time-limited URL. Download and save the file promptlyThe following uses are prohibited:
See the Terms of Service for details.
The Advanced TTS API (Beta) provides 98.0% monthly uptime as its SLA. Both job creation (POST) and polling (GET) requests are counted, and failed speech generation (a polled job whose status is failed) does not count as downtime.
How uptime is measured, what is excluded, and what happens if the SLA is missed are common to both APIs, and the SLA section of the Standard TTS API is the authoritative definition. The SLA covers uptime only, not response time.
Breaking changes (removing fields, changing types or meaning, retiring endpoints, and so on) are announced in this section 30 days before they take effect. Notable backward-compatible changes such as new fields or voices are also recorded in this section after release (no advance notice). Urgent security fixes are an exception.
| Date | Change |
|---|---|
| 2026-10-02 | Launched the Advanced TTS API as an official API (Beta) and set an SLA (98.0% monthly uptime). Started the rule of announcing breaking changes in this section 30 days in advance |
If you run into a bug or unexpected behavior, try the following first:
{"text":"This is a connectivity test.","voice":"Ruby"} (text must be at least 10 characters) returns 202If that does not solve it, contact us through the contact form. Including the following information helps us reproduce the issue:
About specification changes: API specification changes are recorded in the Changelog. Breaking changes are announced 30 days before they take effect.