AgileTTS — API

The API contract: two layers, one service. The front end is the OpenAI-compatible interface with keys (for the products, when AgileTTS is mature); the engine is the resident synthesizer, reachable only from the internal WireGuard network.

The two layers

layerversionwherestatus
engineAGILETTS-v0.9 + attack gateinternal WireGuard networkproduction, resident (WARM), 0.6B model
front endAGILETTS-FRONTE-v0.6https://api.agiletts.agile.softwaredelivered, not yet live

Authentication (front end)

Key in the Authorization: Bearer ats_… header or X-Api-Key. Each key is for one product and one tenant, with a scope (speech, stream, voices), a per-minute limit and a monthly character quota. Texts are never logged: only requests, characters and times are counted.

Test passe-partout key: for first-time integrators, a test-scoped key with low limits (20/min, 20000 characters/month by default) and a declared expiry (24 h by default), meant for a smoke test before requesting a product key. Once expired it returns 403 KEY_EXPIRED. The front also accepts the shared passe-partout key common to every Agile Software sovereign service, the same one for each. Request it from the internal switchboard.

Front end — OpenAI-compatible routes

Method and pathWhat it does
POST /v1/audio/speechspeech synthesis: model = agiletts (aliases tts-1, tts-1-hd, gpt-4o-mini-tts), input ≤ 4096 characters (over 2000 the front end splits at sentence boundaries and rejoins), voice, response_format = wav | pcm | pcm16k | mp3 | opus | flac | aac, speed 0.5–2.0, stream: true for chunked audio (pcm/pcm16k only)
GET /v1/models, GET /v1/models/{id}the available models, with created, owned_by, aliases; /v1/models/tts-1 returns the object with alias_di: "agiletts"
GET /v1/voicesthe voices the key can use: id, language, lexicon, rate, fallback policy — never the reference nor external ids
GET /v1/usageyour key's usage: requests, characters, milliseconds, audio produced, errors, key state and rotations — never the texts
GET /v1/admin/keys, POST /v1/admin/keys, GET /v1/admin/keys/{id}, POST /v1/admin/keys/{id}/rotate, POST /v1/admin/keys/{id}/revokethe per-company console: a customer lists, creates, rotates and revokes its own keys on its own, with a console credential (never a product key). The key in clear is shown only on creation; another company's keys answer exactly like a key that does not exist; every operation, and every refusal, stays written down Since 30/9 the company is named and not guessed: whoever asks for the keys or the people of another company gets a refusal, where before they got their own with nobody telling them.
GET /v1/admin/utenti, POST /v1/admin/utenti, GET /v1/admin/utenti/{id}, POST /v1/admin/utenti/{id}/sospendi, POST /v1/admin/utenti/{id}/riattiva, POST /v1/admin/utenti/{id}/ruoloa company's people, from the console: the administrator lists, invites, suspends, reactivates and re-roles their own users on their own. Whoever suspends themselves, and anything that would leave the company without active administrators, is told no; people of another company answer like a person who does not exist; every operation, and every refusal, stays written down Since 30/9 the company is named and not guessed: whoever asks for the keys or the people of another company gets a refusal, where before they got their own with nobody telling them.
GET /v1/admin/operazionithe record of who did what inside the company, read by the company: who issued a key and when, who revoked it, who invited whom, and the refused attempts too. It can be filtered by person, by action and by time range, and it is read in pages. A plain user does not see it; rows of another company are never there, and whoever asks for them gets a refusal instead of their own company without being told. Reading it leaves its own row as well: looking at a company’s data is an access to its data.
GET /v1/admin/tenants/{id}/piano, PUT /v1/admin/tenants/{id}/pianothe plan agreed with the company — how many requests per minute, how many characters per month — written down where the service can enforce it. Only whoever administers the whole service writes it: a company administrator who could raise it themselves would have a ceiling in name only. They read it, together with the month’s consumption, and that number is the same one the service refuses on. A company with no plan answers «no plan», not «unknown company»; another company and one that does not exist answer the same way
GET /v1/admin/licenze, POST /v1/admin/licenze, GET /v1/admin/licenze/{id}, POST /v1/admin/licenze/{id}/sospendi, POST /v1/admin/licenze/{id}/riattiva, POST /v1/admin/licenze/{id}/rinnova, POST /v1/admin/licenze/{id}/revoca, POST /v1/admin/licenze/{id}/chiavi, DELETE /v1/admin/licenze/{id}/chiavi, GET /v1/admin/licenze/{id}/uso
(the same routes without the version prefix too, the form shared by every Agile service: POST /admin/licenze, GET /admin/licenze, GET /admin/licenze/{id}, POST /admin/licenze/{id}/sospendi, POST /admin/licenze/{id}/riattiva, POST /admin/licenze/{id}/rinnova, POST /admin/licenze/{id}/revoca, POST /admin/licenze/{id}/chiavi, DELETE /admin/licenze/{id}/chiavi, GET /admin/licenze/{id}/uso)
the licence: a company's (or a person's) right to use AgileTTS, with plan, scopes, monthly quota, per-minute rate, start and expiry. Keys live under a licence: suspend the licence and all its keys go quiet together, reactivate it and they resume — with nothing redistributed to the products; rotating a key does not touch the licence. The state that counts is the real one: a licence whose expiry was yesterday reads as expired even if the table says active, and its keys get a refusal that says which licence and why. An expired licence is renewed (not reactivated); a revoked one does not come back: another is issued. Usage is measured per licence and it is the line that goes on the invoice, revoked and rotated keys included. The starting numbers of the prova plan come from the service's measured capacity: 50000 caratteri per month, 10 requests per minute, 30 days (one per cent of the declared service capacity, 5.6 million characters per month per card). standard and enterprise are bespoke and have no starting numbers: they are written at issue time, and asking for one without numbers is a refusal that says so. The lifecycle belongs to whoever administers the whole service; the company administrator reads their licences and issues keys on them within the limits. A licence belonging to another company answers exactly like one that does not exist
POST /v1/richieste-licenza (public, no key), GET /v1/admin/richieste
(without the version prefix too: POST /richieste-licenza, GET /admin/richieste)
the «request a licence» form of the public page: company, contact person, email, plan and notes. It does not send an email and it does not swallow: the request is recorded with a number and the page says it («request registered no. 7»), a round on the box carries it to whoever must read it and marks it notified only if the notice really went out. Prices are in no answer at all: they are on request. Two guards, because the route is public: a per-address request limit and a hidden field a person never sees — if it arrives filled in a robot filled it, and the request is refused saying so, not accepted for nothing
POST /tts, POST /tts-streamhistorical pass-through: same body as the engine, with a key → injecting into the products means changing the address and adding the header
GET /healthstate of the front end and the engine: model_loaded, caldo, coda, trim_attacco, normalizza, formati, testi_lunghi, alias_modelli, politiche

Engine — internal routes

Internal WireGuard network only, no authentication. JSON in, audio out, one synthesis at a time. The request body is the same as the front end's historical pass-through.

Method and pathWhat it does
GET /healthengine state: version, model_loaded, warm, free VRAM, trim_attacco, loaded voices, counters (requests, synth_total_s, identity_retries, …)
GET /voicesthe voices with text reference, lexicon, rate (atempo_target_wpm) and per-voice parameters
POST /ttsblock synthesis: text (required, ≤ 2000 characters), voice (default patrizia), format = wav | wav16k | pcm16k | mp3 | mp3_16k, apply_lexicon, atempo, trim_attacco. Response: audio body + X-AgileTTS-Meta header
POST /tts-streamchunked streaming: same fields, format = pcm24k (default) | pcm16k; PCM s16le mono body in chunks until it ends
POST /warm, POST /admin/unloadloads the model (/warm) or unloads it (/admin/unload, regia only); in production WARM reloads it at the next request

Audio formats

Errors

HTTPcodemeaning
401KEY_MISSING, KEY_INVALIDkey missing or unknown
403KEY_REVOKED, KEY_ROTATED, KEY_EXPIRED, SCOPE_NOT_ALLOWED, PLAN_EXCEEDEDkey revoked, rotated (grace ended), expired (test key) or without the requested permission, or a key that on its own would promise more than the company plan
404UNKNOWN_VOICE, UNKNOWN_MODEL, TENANT_NOT_FOUNDvoice or model that does not exist, or (only for whoever administers the whole service) a plan written on a company that does not exist
400BAD_REQUEST, UNSUPPORTED_FORMATinvalid body or unsupported format
409ALREADY_EXISTS, SELF_SUSPEND_FORBIDDEN, LAST_ADMINon the per-company console only: person already invited, an administrator suspending themselves, or an operation that would leave the company without active administrators
413TEXT_TOO_LONGover 4096 characters on the front end (over 2000 on the engine)
429RATE_LIMIT, QUOTA_EXCEEDED, TENANT_RATE_LIMIT, TENANT_QUOTA_EXCEEDEDtoo many requests per minute or monthly quota exhausted: of the single key, or of the whole company on the sum of its keys (with Retry-After)
503BUSY, ENGINE_DOWN, SHUTTING_DOWNengine busy (with Retry-After), VRAM guard: unavailable, or the front is stopping for a redeploy (with Retry-After: whoever is already speaking finishes, whoever arrives retries in a few seconds)
502ENGINE_ERRORinternal engine error
500synth_failed, stream_failedsynthesis failed (block) or stream failed (only before the first bytes)
501pcm16k_unavailable, streaming_not_availableformat or streaming unavailable on the engine

Front end errors use the OpenAI shape {"error":{"code","message","type"}}. On 5xx the front end never falls back by itself: it tells the client what it can do (X-AgileTTS-Ripiego and error.ripiego), the decision stays with the caller.

Response headers

Example

curl https://api.agiletts.agile.software/v1/audio/speech \
  -H "Authorization: Bearer $AGILETTS_KEY" -H "Content-Type: application/json" \
  -d '{"model":"agiletts","input":"Buongiorno, come posso aiutarla?","voice":"serena","response_format":"wav"}' \
  -o saluto.wav

With the official OpenAI client you just change base_url and the key: model: "agiletts" (or "tts-1") and response_format: "opus" work without any code change.

Clients for products

Besides the official OpenAI client (which talks to the front end without changes), the repo ships an example client in pure standard library, client/agiletts_client.py. For those who prefer the command line there is client/agiletts-cli.py, the command-line client in pure standard library that wraps client/agiletts_client.py: block or streaming synthesis to a file (-o), models/model/voices/usage, key from --api-key or AGILETTS_API_KEY, errors on stderr with a non-zero exit code. The step-by-step guide to move a product onto the front end — URL, key, format, timeout, 503 BUSY and fallback, the "go" and the rollback one at a time — is in docs/ADOZIONE.md.