AgileTTS — News

One entry per date: what changed, who it is for and how to adopt it. Entries marked «delivered, not live» become «live» with the date of the regia's «go». The full contract is in API.

30/9/2026 — The public address serves the service, not its control room delivered, not live

We are about to open a public address for AgileTTS, so that a product already in production can use our voice instead of its previous supplier. The piece in front of the service, though, was letting everything the service can do through on that address: not only speech, but governance too — the console a company uses to manage its keys and its people, the licences, and the licence requests received, which belong to whoever governs. Before saying "ready" we counted it one route at a time, reading the list from the service's code and not from a hand-written list — a hand-written list always forgets the thing born yesterday. The count: of the twenty governance paths, zero were closed. Now twenty out of twenty, and the product still passes through in full, ten out of ten, public page form included. The closure is by family and not by list: whatever is born tomorrow under governance is born closed, instead of being born open and waiting for someone to notice. And the service, for its part, listens to itself only: measured on the computer hosting it, it does not even answer from the group's private network, so whoever arrives from the public address necessarily passes through our filter. That detail came from an implicit value, and an implicit value flips over in silence: it is now written down where the service is started.

Who it is for: the product about to adopt us, which gets the voice without getting our control room along with it; whoever opens the address, who now knows which paths go out and with what count; and administrators, who know where one comes in from, and know we measured it for them instead of deducing it.

How to adopt it: nothing to change for whoever uses the service with their own key — on the public address speech is identical. Whoever administers (console, licences, requests received) comes in from inside the computer hosting the service: on the public address those paths answer "does not exist", and that is not a fault, it is the perimeter. One detail today's measurement corrected: since the service listens to itself only, today the console is not reachable from the group's private network either — only from inside. How it should be reached is a choice for whoever governs, and we have put the question. The public address is not live yet: what is closed here is the configuration, and the check from the seat of whoever will use it will happen once the name is really online.

30/9/2026 — A licence belongs to one company, and so does a key delivered, not live

A licence belongs to one company; a key belongs to one company. As long as the two match, the licence above the key is a contract. When they do not, the victim's licence pays for a stranger: its per-minute ceiling, its monthly quota and the row that ends up on the invoice. The question reached us from the sister service and we measured it on ours instead of taking it for granted: neither of the two places that write the key→licence link compared the two companies. On a throwaway store: a thousand characters out of a thousand of the victim's quota consumable by someone who is not its customer, the stranger's key listed in another company's licence detail, and the victim's per-minute seat taken by him. Now zero. We closed two doors, not one: moving a key issued under a licence into another company is refused, and the refusal names the licence; and on every request a key whose licence belongs to another company gets a no that says which licence, served before the per-minute ceiling is counted — otherwise the victim's seat would already be taken by the time the no arrives. Keys without a licence, the older ones, are linked as always: closing that case too would have been worse than the defect.

Who it is for: every company holding a licence, because its ceiling, its quota and its invoice can no longer be consumed by a key that is not its own; and for administrators, who on trying to move a key that sits under a licence read a no with the licence's name, instead of finding out from a wrong invoice.

How to adopt it: nothing to change for whoever uses the service with their own key. If a key must pass to another company, revoke it under its licence and issue a new one under the right licence: moving the row is no longer possible, and that is exactly the point. Whoever receives this no should not change key: it is a stored row to be fixed, and it is reported to the administrators.

30/9/2026 — The public page's form was asking the browser's permission of nobody delivered, not live

The public page sits at one address and the service at another: to a browser these are two different worlds, and before sending a form from one to the other the browser asks for permission. The service did not answer that question at all, and it did not put on its real answers the line that lets a browser read them either: the «request a licence» form and the live service status — two of the six elements of the page delivered this morning — did not work from a browser, while answering command-line tools perfectly well. Measured from the browser's seat: two elements out of eight, now eight out of eight. Permission is granted by name, never to anyone, and only on the three things the page actually asks for: the routes that need a key stay closed to the browser, because a key belongs on the server of whoever owns it, not inside a page. An address outside the list gets a refusal that names it, instead of the mute error a browser usually shows. In the same place the form's ceiling is fixed: it counted the address of the connection, but in front of the service there is a proxy on the same machine, so that address is always the same and the ceiling would have been one for the whole Internet — three requests a minute in total, with the first robot shutting the form in every real customer's face. It now counts the customer, and believes whoever declares it only when the request comes from our own proxy.

Who it is for: anyone opening the public page, because the form and the service status now really work from a browser and not only from technical tools; and whoever receives licence requests, who would otherwise have seen zero of them arrive from a page that looked perfectly fine.

How to adopt it: nothing to change for whoever uses the service with their own key. Whoever serves the page from another name must declare it in the service's configuration, otherwise the form goes quiet in the browser while still answering command-line tools — and that is exactly the case where the refusal says which name is missing. Whoever puts the service behind a proxy other than ours must make sure that proxy declares the customer's address, or the form's ceiling will go back to counting the proxy.

30/9/2026 — Licences: the right to use AgileTTS now has a name, a state and a bill delivered, not live

Until today whoever used AgileTTS had a key, and the key was everything: the permission, the ceiling and the contract all at once. What sat in between a company and its keys was missing: the licence. We measured it before writing a single line, by asking the real service for the ten routes the standard shared by every Agile service expects: zero answered. From today a licence says who it belongs to, which plan it is, which permissions it covers, how much can be spoken in a month, how many requests per minute, from when and until when. The keys live under it: suspend the licence and all its keys go quiet together, reactivate it and they resume — with no product having to receive a new key. The state that counts is the real one: a licence whose expiry was yesterday is expired even if the archive says «active», and its key gets a refusal that says which licence and why, not a generic «invalid key». An expired licence is renewed; a revoked one does not come back: another is issued, and the list shows that it is another one. Consumption is counted per licence and it is the line that will end up on the invoice: switched-off keys stay inside it, because what has been spoken has been spoken. And the plans' starting numbers are not picked by eye: they come from the service's measured capacity (how many characters per month one card sustains, with headroom for the spike), and the trial is one per cent of that capacity. Prices stay on request.

Who it is for: every company integrating with us, because «from when to when, for how much, with which permissions» is now written in one place and read from the console, instead of living in the memory of whoever issued the key; whoever runs the service, who suspends or reactivates a whole company without touching its keys one by one; and whoever invoices, who will get one line per licence instead of a sum to redo by hand.

How to adopt it: nothing to change for whoever uses the service with the key they have today — keys born before licences keep working identically, and switching them off because «they have no licence» would be a failure dressed up as a rule. Administrators find licences in the console, with their administrator credential: creating, suspending, renewing and revoking belongs to whoever runs the service, while the company administrator reads their own and issues keys on them within the licence's limits.

30/9/2026 — The public page, with the plans and the form to request a licence delivered, not live

The AgileTTS site was documentation: what the API is called, which routes exist, which errors. For whoever has to decide whether to use it there was nothing: not who it is for, not which plans exist, not how to request a licence, not whether the service is up right now. Of the seven elements a public page must have, ours had two; now it has six, in Italian and in English: who it is for, the three plans with what they include and the price on request (zero figures on the page: we do not decide prices), the link to the documentation, the live service status — the page asks the service and says so, and if it does not answer it writes «status unavailable» instead of staying blank — and the «request a licence» form, which does not send an email and does not swallow: it records the request with a number and tells you, and a round carries it to whoever must read it, marking it notified only if the notice really went out. The seventh element — the comparison with the leader using real numbers — is missing, and the page declares it: our six numbers are there, each with its source (including the awkward one), but of the leader's numbers we have none measured, because the key needed to test it has been blocked since 21/9. An invented number on a public page is worse than a missing one.

Who it is for: whoever evaluates AgileTTS — an internal product, a customer, whoever proposes it — and until today had to ask everything by voice; whoever integrates with us, who sees from the same page whether the service is up; and whoever receives the requests, which now arrive with a number instead of by word of mouth.

How to adopt it: nothing to change for API users. The public names of the page and of the API are created by whoever runs the network and are not active yet: until they are, the page is viewed from the artefacts. Whoever serves it from another address must change the form's line, which calls the service to record the request.

30/9/2026 — Asking for somebody else's company is a no, not your own in silence delivered, not live

In the console you can say which company you are working on, because whoever administers the service works on all of them. For an ordinary company that name was ignored: whoever asked for another company's keys or people got their own, with nobody telling them — measured, 2 cases out of 2. Whoever then copies that list into a report believing it belongs to the company they asked for never notices. From today, if the company is not yours the answer is a refusal; if it is yours, it passes as before. And the refusal leaves its row in the record of whoever asked, not in that of the named company: otherwise a stranger could write rows into the record of a company not their own.

Who it is for: every company integrating with us, because a list read by mistake is no longer indistinguishable from one read by right; and whoever writes the integration, because a silently ignored parameter is a mistake you cannot see until it does damage.

How to adopt it: nothing to change for whoever uses the service to give their products a voice. Administrators who named their own company carry on identically; those who named another one expecting their own things now get a no.

30/9/2026 — The record of who did what is read by the company delivered, not live

Since 29/9 a company manages its own keys and people, and every action leaves its row in a record: who issued a key and when, who invited whom, who suspended whom, and the refused attempts too. That record, though, only we could read, from our own box. Measured: a working day writes 12 rows and the company's administrator could read zero of them — the five questions one asks the day after were answered 0 out of 5 from the console and 5 out of 5 by us. From today the company reads those rows itself, most recent first, filtering by person, by action or by period: 12 out of 12 and 5 questions out of 5. Rows of another company are never there, and whoever asks for them is told no instead of being handed their own without a word; the name of whoever acted appears only if they are a person of that company; and looking at that record is itself an access to the company's data, so it leaves its own row.

Who it is for: every company that has to answer "who did this, and when" — for an internal review, a security audit, a colleague — without asking us anything; and for us, who stop standing in the middle of every question.

How to adopt it: nothing to change for whoever uses the service to give their products a voice: no call changes and no field disappears. Whoever administers finds the record in the console, with their administrator credential; a plain user does not see it.

30/9/2026 — A redeploy no longer throws away whoever is speaking delivered, not live

When we update AgileTTS we restart it, and restarting means telling the program "stop". Until today the program stopped in that instant: whoever was speaking at that moment lost the sentence halfway, and did not even get an explanation — only a network error, which inside a product is the branch that falls back to an American vendor. That is: a silent fallback, caused by us, on every update. We measured it before touching the code, with the real service, the queue full as in production and the "stop" halfway through a sentence: 0 calls out of 3 got their audio, 3 out of 3 were left without audio and without an explanation, and no line was left in the log — so not even we could notice. From today a service that stops does three things: it finishes saying what it had started (a sentence lasts a few seconds, and the card had already done the work); to whoever arrives at that moment it answers "I am stopping, retry in a few seconds" with a code and a number, and it does not put that error on their account, because it is not their fault; and it declares it on its status page, so whatever routes the traffic takes us out of rotation before the address goes silent and monitoring does not raise an alarm for a deliberate update. Same measurement after: 3 out of 3 get their audio, nobody is left without an explanation, the log carries every line and the address stays silent for a quarter of a second. Two things were found by the bench and not by reasoning: with the service idle the exit cost half a second too much, given away on every restart; and the time we wait is not a number picked by eye, it is how many sentences the card can have in its mouth times the length of one sentence, because the card says one at a time. That is why, when that time is not enough, the service writes it in the journal with the number of sentences cut off: it is not for a client to discover.

Who it is for: the 081 phone line first, which calls us while somebody is on the line; anyone integrating with us, because a deliberate stop now has a name and a number instead of looking like a failure; and ourselves, because it is the precondition for updating the service during working hours instead of waiting for the night.

How to adopt it: nothing to change. There is a new code for the "I am stopping" refusal, and anyone already handling temporary refusals with a "retry after" does not have to write a line. Whoever installs must take two new lines of the system configuration, because without them the operating system would kill exactly whoever we are waiting for — and a test turns red if the two numbers stop agreeing. One thing we say because it stays open: a very long text (the maximum we accept) needs more than nine minutes of card time, and no waiting window will ever cover it: sent in the instant of an update it will be cut off, and the journal will say so.

30/9/2026 — A refusal you can retry says when, with the true number delivered, not live

Some refusals are not failures: the per-minute ceiling on calls, the month's quota exhausted, the machine full, the machine down. For all of them the caller's question is the same — when may I retry? — and if we answer badly the one who pays is the polite client, the one who listens to us. Measured before touching the code, with the real AgileTTS and real calls: we said 60 seconds where the minute reopened in 35 (twenty-five seconds of waiting given away), we said 60 seconds to whoever had used up the month's quota, which reopened in 85534 — a thousand times as much, and a polite client knocked once a minute for weeks — and we said 2 seconds in front of requests we serve in 6.5, with twelve clients refused together all coming back in the same second. From today the number is the true one: for the window ceilings it is the instant the window reopens, for the full machine it is what is left of the request in progress, measured on what AgileTTS has just served. Afterwards: one knock instead of three and served after 6.6 seconds, i.e. when the seat actually freed; and of the twelve refused together, at most five come back in the same second instead of nine.

Who it is for: every product that calls us and then does not have to invent a waiting policy of its own; and us, because every knock on a closed door is work the card does for nothing.

How to adopt it: nothing to change, no field disappears. The number is read by the HTTP libraries themselves; whoever wants the exact instant to the second finds it in the body of the refusal. The number told on the month's quota is capped at one hour on purpose: a plan can be raised at any moment, and saying "come back tomorrow" would keep a product out even after we have widened its quota.

29/9/2026 — The ceiling belongs to the company, not to a single key delivered, not live

With a company we agree a plan — how many calls per minute, how much text per month — and from today we give it a console to make its own keys. The two together had a hole, and nobody had ever measured it: the per-minute limit and the quota lived on the single key, and from the console an administrator makes as many keys as they like — one per product, which is exactly what we ask them to do. We measured it before writing a line of code, with the real service, the real console and real calls: with the plan "six requests per minute, six texts per month" and three keys made in good faith, eighteen went through per minute and three times the month's text; a key with "no limit" was accepted under that plan; and the plan could not even be written down anywhere — it was agreed verbally, and the service did not know it. From today the plan belongs to the company and bites in two places: when issuing a key that on its own would promise more than the plan (refused, "no limit" included), and on every request, on the sum of all the company's keys. Same measure afterwards: one to one, the plan respected. Four choices deserve a line. (1) Only we write the plan: if the company's administrator could raise it themselves it would be a ceiling in name only. They read it, with the month's consumption, and that number is the same one the service refuses on — otherwise console and service would say two different things about the same company. (2) Issuing looks at the single key, every request looks at the sum: refusing on the sum right away would forbid the second key to whoever has two products, which is the very use the console exists for. (3) A key rotation is never blocked by the plan: it copies the limits of a key that already exists and promises nothing new, and blocking it would deny a security rotation to a company whose plan was lowered afterwards. (4) The refusal says when to retry, with the exact number: the per-minute ceiling has a precise instant when it reopens and the service knows it, so it says the seconds left instead of a round minute. One defect was found by the measure and not by reasoning: with the two checks in a row, a request stopped by the company's ceiling had already consumed a slot of the key's rate, and right afterwards the client was told "too many requests on this key" — the wrong limit, the one an administrator would raise without fixing anything. Now a slot is taken in both ceilings, or in neither.

Who it is for: the customer companies, which at last see their own plan written down and know where a refusal comes from; and us, because without this the console we deliver today handed out the way to multiply the agreed plan by the number of keys, without anybody lying. For whoever uses AgileTTS to speak (core/081, adp, nis2 voce, vocea, dfm) nothing changes: their keys are issued by us from the box, belong to no console company and no plan can concern them — and this we measured, we did not deduce it.

How to adopt it: nothing to change, no field disappears, no key changes by a byte. A company with no plan written behaves exactly as before: the plan is written when it is agreed, there is no default. New fields: the key's usage page now also carries the company plan and how much of it has been consumed in the month, so a product that receives a refusal sees where it comes from without a console credential it does not have. One thing we say because it stays open: across the turn of the minute twice the requests still go through in a few seconds, because the window is fixed — it holds for the key as for the company, and it is a change of its own that we will make and declare.

29/9/2026 — Two machines of different speed: the call goes to the one that answers first delivered, not live

With two machines behind AgileTTS, a call went to the one with the least work in progress, and round by round among equals. But calls arrive one at a time: the machines are almost always idle, "the least busy one" is a nil-nil tie, and the turn kept deciding. We measured it before writing a line of code, with the real service, two machines of different speed and real clients: the slow machine took half of the calls and the client waited 0.65 seconds on average against 0.15 for the fast machine alone — four times the best achievable, paid for with the fast machine switched on and idle. We already had the number we needed: the service measures the speed of every machine on what it has just spoken (it has used it since 28/9 to tell whether the conversation cap is still honest). From today it also uses it to choose: machines are ordered by what the call would cost on each — the work ahead of it, times how fast that machine is — and the first one with a free seat is taken. Same measurement afterwards: zero per cent to the slow one, and the client waits what they would wait with the fast machine alone. Three choices deserve a line. (1) Speed is learned by serving, not declared: nobody writes "this is the slow one", because a hand-written number goes stale with the first model change and no one notices; in the measurement ten calls were enough. (2) A machine just added, which we have not measured yet, counts as the fastest one and not as the worst: giving it a cautious weight would mean sending it no calls, hence never measuring it, hence leaving it forever with the weight of the unknown. (3) It is a preference, not a ban: when the preferred machine is busy the call goes to the other one and is served — with four simultaneous clients, four out of four still get in, capacity does not change by one seat — and the voice always comes first: a voice that lives only on the slow machine goes to the slow machine.

Who it is for: anyone who tomorrow puts a second, different card next to today's one, which is the normal case — cards are bought when they can be found, not identical to the previous one. Without this, half the clients ended up on the slower machine and paid the difference in waiting while the fast one sat idle: one paid for a card in order to make the service worse for half the calls. For those who use AgileTTS to speak (core/081, adp, nis2 voce, vocea, dfm) nothing changes: nobody has a second machine today, and with a single one there is nothing to dispatch.

How to adopt it: nothing to change, no field disappears and no new setting is mandatory. Whoever runs more than one machine need not declare which one is the fastest: the service measures it itself. The status page says whether the choice looks at speed and, for each machine, how many times it costs more than the fastest one. Whoever does not want it can go back, with one setting, to the previous dispatching, which looks only at the work in progress.

29/9/2026 — The second machine adds seats, not just an address delivered, not live

Since this morning AgileTTS can stand in front of several engines. What we had not measured is the very reason one adds an engine: capacity. We measured it before writing a line of code, with the real service and real clients: with the conversation cap at two and four simultaneous clients, one machine let two in; two machines let still two in — the second added zero seats. And the client of a voice living only on the second machine got a "try again shortly" while that machine was idle at zero requests: it was paying the queue of a card that was not serving it. The reason: the cap belonged to the service, i.e. to the process relaying bytes, while what fills up is the card. From today the cap belongs to one machine: two machines are two queues, and above them stays a service-wide net that can be set separately. Same measurement afterwards: four clients out of four, and the second machine's client served. Three choices deserve a line. (1) The machine is chosen and the seat is taken at the same instant: counting them at two different moments means two requests see the same seat free and both take it. (2) Dispatch looks at the load and not only at the turn: round-robin distributes turns, not work, and kept sending requests to the full machine while the other stood idle; at equal load it stays the previous rotation, so with idle machines — the normal case — nothing changes. (3) The refusal now says which of the two caps spoke and names the full machine, because the two remedies are different: add a machine, or widen the service. The reason we had written for keeping a single cap — "we do not yet know which machine will serve the request" — was contradicted by the code: the machine was already chosen first.

Who it is for: us and the owner, because "adding an engine changes nothing for the customers" only holds if adding it adds capacity; and anyone who tomorrow puts a second card next to today's one, who would otherwise pay for it only to leave it idle. For whoever uses AgileTTS to speak (core/081, adp, nis2 voce, vocea, dfm) nothing changes: nobody has a second machine today, and with a single one the cap is exactly the previous number.

How to adopt it: nothing to change, and no field disappears. Whoever runs more than one machine need not touch the number they had already written: only its meaning changes, from "the whole service" to "one machine", which is what it always meant. Whoever wants the service tighter than the machines can set a separate net. The status page says, for every machine, its name, how many seats it has taken and how many it has.

29/9/2026 — A company administers its own people from the console, and no company is left without an administrator delivered, not live

Since this morning the front's records know who a company's people are, but nobody could reach them from outside: to add a colleague to the panel, remove their access the day they leave, or appoint a second administrator, you still had to write to us. Before writing a line of code we measured how much a company could do on its own: of the six operations on its own people (see them, invite one, read one, suspend, reactivate, change a role) zero out of six were possible — the call simply did not exist, which is a different thing from "it answers empty". From today it is six out of six, and each one stays written down with who did it and when. Three things deserve a line. (1) No company is left without an administrator. The same measurement found two open doors that had nothing to do with the new calls: a company's last administrator could suspend themselves and could demote themselves to a plain user, and in both cases the company was left with nobody who could administer it, not even to reopen the door. They are now two declared refusals, and the check sits on the data and not only in the call: closing the suspension alone would have left the other door open. Where a second administrator exists, on the other hand, the demotion goes through: a check that blocks everything is of no use to anyone. (2) The service-wide administrator role is neither assigned nor invited from here, not even at our own request; and inviting someone already holding a role goes through the same check as assigning it — that would be the service door to the same permission. (3) People of another company answer exactly like a person who does not exist: same answer, same text, so nobody can discover who works elsewhere by trying identifiers one by one. Two defects were found by the test and not by reasoning: inviting the same person twice looked like a service fault (now it is a clear refusal), and whoever administers all companies, if they did not say which one, silently got the list of one of them (now they are asked to say).

Who it is for: the customer companies, which no longer have to wait for us to manage their own people; and for us, staying only where it is really needed — creating a company and appointing its first administrator. For whoever uses AgileTTS to speak (core/081, adp, nis2 voce, vocea, dfm) nothing changes: no new call on the synthesis paths, no new field, no key touched.

How to adopt it: to speak with AgileTTS nothing changes, and whoever has no console credential sees nothing new. To administer your own people you need the administrator's credential, which today we issue, and six calls: the list (which also says how many active administrators the company has), the invitation, the read, the suspension, the reactivation and the role change. No records to update: they are this morning's. The console is not live yet: it comes up with the next deployment, and its web page is the next step.

29/9/2026 — A company issues its own keys from the console, not from root delivered, not live

For a new key, or to rotate one, until now you had to write to us: keys were issued only from the command line, on the service's machine. Before writing a line of code we measured how much a company could do on its own: of the four operations on its own keys (see them, create one, rotate it, revoke it) zero out of four were possible, and zero out of four left a record of who had done them. From today it is four out of four, and each one stays on record. Three choices deserve a line. The console does not open with the product key: it takes a different credential, because a key stolen from a phone must not be worth an administrator's password — and the console credential, conversely, cannot make the service speak. Another company's key answers exactly like a key that does not exist: same answer, same text, so from the outside nobody can discover which keys others hold by trying identifiers one by one. And permissions are checked before the archive is touched: refusals stay on record too, with who tried and when.

Who it is for: whoever today has to wait for us for a new key or a rotation, and us, who stay only where we are really needed — creating a company and appointing its first administrator.

How to adopt it: nothing changes to talk to AgileTTS — the usual calls are untouched, the archives update themselves and no existing key changes. To use the console you need a credential, which today we issue for the company's administrator. The console is not live yet: it comes up with the next deployment, and its web page is the next step.

29/9/2026 — Who invited whom, and when: the foundations of the per-company console delivered, not live

The front's store knew, for each key, an owner and a company — two free-text fields. It knew nothing about people: who rotated a key, who invited whom, who was suspended and when, could not even be asked. Measured before writing any code, on the 8 questions a multi-company console must be able to put to the store (which companies exist, which users a company has, what role a person holds, who is suspended, who invited whom and when, who changed a role, which company a key belongs to by identity, all of a company's operations in time order): 0 out of 8 could even be formulated — not "they answer empty", the question did not exist — and 4 state-changing operations out of 4 were silent. From today there are companies, users with role and state, and the operations log; and a key ties to a company by identity instead of by the previous free text. After the same measurement: 8 questions out of 8 answered and 12 management operations out of 12 leaving their log row. Three choices that go beyond "adding the tables". (1) The log row sits in the same transaction as the change: if the trace cannot be written, the change does not stay — a store must never end up in a state nobody can explain, and remembering to call the recorder is not enough. (2) The log never carries a key's cleartext: a detail containing it is refused, not silently scrubbed — scrubbing would let the mistake repeat elsewhere. (3) The superadmin role is assigned by no function, not even by a company administrator: the guard sits on the data and not only in the route that will arrive on top. The migration is additive and idempotent: on a store already in use everything adds itself, no existing key changes by a single byte and old keys keep working. How to adopt it: nothing in the clients and nothing to do for whoever administers it — the migration runs by itself on the first connection, as the rotation columns already do. The self-service console is not there yet: these are the foundations, the routes are the next step.

Who it is for: the owner and regia, because this is point A of the "complete service" mandate (a company has its own users with roles, and one company's data is never visible to another) and the base the self-service routes will rest on; and every customer company that today, to get a key issued or rotated, must ask a human with administrative access to the front's box. For the already integrated products (core/081, adp, nis2 voce, vocea, dfm) nothing changes: no new route, no new field in the responses, no key touched.

29/9/2026 — One front in front of N engines: per-voice dispatch and parallel probes delivered, not live

Until today the front pointed at one engine: adding capacity meant a second front, i.e. a second contract, a second key store and a second usage metering. From today, with AGILETTS_MOTORI set (N comma-separated addresses), a single front stands in front of N engines and dispatches every request round-robin across the healthy ones. Three choices, all measured. (1) Dispatch looks at the voice: an engine only receives the voices it declares. Without that filter two engines with different catalogs would return 404 UNKNOWN_VOICE intermittently for the same voice — the request succeeds or fails depending on where the round-robin lands, which is the worst way to break. (2) The hot path never probes: probes run in parallel and sit in a 5-second cache. The obvious way — asking every engine, in sequence, at request time — has a cost you don't see while everyone answers, and you pay it all at once when an engine goes mute (process alive but stuck: it costs the whole timeout, whereas a dead process refuses immediately). Measured with real mute engines: probes in sequence 2.06 s with one mute and 8.01 s with four, in parallel 2.01 s in both cases; the front's /health stays flat at 2.01 s from 1 to 4 engines; on the hot path 6 syntheses cost 0.36 s with the cache against 12.39 s probing on every request (per-request p95 72 ms against 2071 ms) — in front of a product ttfb p95 of 0.88 s, probing every time would have tripled the response time to keep up with engines nobody was using. (3) An engine that dies is marked down at once: one request falls on the dying engine, not all of them, and the periodic probe puts it back in service by itself when it recovers — no blacklist nobody knows how to clear afterwards. The front does not retry the fallen request on another engine: that would be a second unrequested synthesis, and it is a decision about the contract that the box does not take by itself. A voice living on an engine that is now down returns 503 ENGINE_DOWN (it will come back), a voice no engine declares returns 404 UNKNOWN_VOICE: they are two different things and the client must be able to tell them apart. GET /v1/voices is the union of the healthy engines' catalogs; /health carries the motori block instead of motore. The queue cap stays aggregate on the front, and the rtf is measured per engine — one window of samples for each address, exposed in motori.dettaglio[].rtf_motore and in coda_onesta.rtf_per_motore — with the verdict on the worst p95 among the healthy engines (it is the slowest engine that decides whether accepting N conversations is still honest); until at least one healthy engine has enough samples the front says "I don't know" and does not declare the verdict.

Who it is for: regia and the owner, because it is point 5 of the "complete service" mandate and it is what it takes to put a second card next to the first without giving the products a second address; and every already-integrated product (core/081, adp, nis2 voce, vocea, dfm), for whom adding capacity behind the front changes nothing: same address, same key, same contract. How to adopt it: nothing in the clients, and nothing to do for those who do not set AGILETTS_MOTORI — without that variable there is not one extra probe or thread, and that is the first thing the test verifies. For two engines behind one front: AGILETTS_MOTORI=http://127.0.0.1:8795,http://secondo-motore:8795 in the front's unit, then GET /health to see motori.sani. Tested by fronte/prova_multi_motore.py (51 checks, 6 mutations with declared blast radius) and measured by banco/tts/capienza_motori.py.

29/9/2026 — The engine patches waiting for the "go" now apply to the real engine, and the GEX44 variable names were documented wrong delivered, not live

On the night of 29/9, inside the "go" window, the guard patch did not apply on the production engine — "anchors missing" — and the launcher rolled back with a pointless restart of the engine at 02:00:42; queue and cap were not even attempted. The three patches had been green for days, but they were tested against the repository seed, and the seed is not the engine: 615 lines against 717, 326 differing lines. The defect, measured on the copy of the live file, has two parts and neither is about what the patches do. (1) The prefix: the seed is renamed and reads AGILE_TTS_*, the live engine reads AGILETTS_*, so an anchor like os.environ.get("AGILE_TTS_MAX_IDENT_RETRY" does not exist on the live one — the first piece of all three fell. (2) A seed-only line: the queue piece that adds the /health block anchored on a line the renaming creates and the real engine does not have, and queue stayed red even once the prefix was solved. Now every anchor is looked up in both forms and the one that appears exactly once wins, and the /health anchor was moved onto the last line of the block, which exists in both. Only the anchors are translated, never the new knobs: those have a public name, and translating them would mean changing the contract silently, at night, because of an anchoring defect. Measurement on the copy of the live file: the three patches apply, the stack compiles, and the three functional tests are green on the live form (13 + 17 + 9 checks), with the bite that requires them red on the unpatched live file. The 30 knobs the engine already read stay AGILETTS_*, the same thirty; the 7 new ones stay AGILE_TTS_* as published. New guard: bin/gex44-notte/prova_seme_vivo.py, with 5 mutations requiring the exact set of fallen checks.

Who it is for: the owner and regia first of all, because the "go" night is theirs — three patches out of four would not have started, and the cost was not zero: a restart of the production engine for a patch that had not changed a single byte. Then whoever configures the engine on the GEX44: the environment variable table in API carried the seed's names (AGILE_TTS_MODEL, AGILE_TTS_WARM, …) as if they were the real engine's, while the production engine reads AGILETTS_* — an AGILE_TTS_* variable on the GEX44 raises no error, it simply has no effect, which is the quietest way to believe you configured something. How to adopt it: nothing in the clients, the HTTP contract does not move and the names of the knobs published by the patches do not change; whoever configures the engine on the GEX44 should use AGILETTS_* for the base engine variables. The patches remain delivered and not applied: they wait for the owner's "go".

29/9/2026 — Test passe-partout key, both self-issued and shared across sovereign services delivered, not live

Anyone integrating AgileTTS for the first time used to have to wait for a product key before running a first smoke test. Now agiletts-keys.py issue-test issues a PASSE-PARTOUT key with test scope (status: "prova"): same shape as a real key, full scope, but with low limits (20/min, 20000 characters/month by default) and a mandatory expiry declared at issue time. Before expiry it behaves exactly like an active key; once expired it returns 403 KEY_EXPIRED — never KEY_REVOKED, because it was not revoked, its declared time simply ran out. On top of that, for the regia's mandate of 29/9 ("every sovereign service ready and complete", with a test passe-partout key accepted by each one), the front also accepts a second key, shared across every Agile Software sovereign service, from the vault (agile__passepartout_test/api_key, "rotates on 29/10"): if the AGILETTS_FRONTE_PASSEPARTOUT env var is set at boot (never in git, never hardcoded), the front installs it itself with the same treatment. Env absent = no shared key installed, unchanged behavior. Tested by fronte/prova_keys.py (26/26) and fronte/prova_auth_fronte.py (two dedicated control mutations).

Who it is for: anyone integrating AgileTTS for the first time who wants a real curl before requesting a product key (core/081, adp, nis2 voce, vocea, dfm); and regia, for whom the shared key is a requirement of the "complete service" mandate common to every sovereign service. How to adopt it: nothing for those who already have a product key; for a smoke test, agiletts-keys.py issue-test on the front's box, or the shared passe-partout key (request it from the internal switchboard).

29/9/2026 — «Texts are never logged» is now measured across the front's whole disk delivered, not live

The guarantee has been in the contract since 15/9, and on 28/9 it was hardened on the service path (err = the exception's type only). But it was measured field by field: a probe reads the request log lines and checks that no field carries the text. That is the right guard on the log, but the log is not the whole disk: what stayed outside the measurement was the process's stdout and stderr — which in production are the systemd journal, i.e. the place where a text would sit for days without anyone having decided it — plus the DB, the usage page, the /health body and temporary files. From today the measurement covers the whole disk, in two clauses. A · no trace of the text: the front runs as a real process (not imported: stdout and stderr are file descriptors, as under systemd) and receives a text carrying a sentinel over 26 paths — block wav/pcm16k/opus, streaming, long text split, the historical /tts and /tts-stream paths used by 081 and the avatar, and every failure and refusal path; then the whole state directory is swept byte by byte with 12 needles: the sentinel, the plaintext keys and the text's sha256 fingerprint, because a fingerprint lets you confirm a guessed text and that is retention too. The sentinel also travels in a query string, because a request line that ends up in the journal would carry the text away together with the path. Measured: zero occurrences across 9 files, with 27 log lines to say the front really had served (a probe through which no text passes proves nothing). B · no abandoned audio: audio is the text spoken and for art. 9 it is worth as much as the text, so the round runs with TMPDIR inside the test directory and demands it be empty at the end — a surviving transcoding temporary file is the client's audio left on disk, and clause A would not see it because that file holds audio, not the token. Tested by fronte/prova_ritenzione.py (three mutations on the new surfaces: a print of the text, 97 occurrences on stdout; log_message no longer silenced, a single occurrence on stderr and that is exactly the point; the temporary file no longer deleted).

Who it is for: vocea first of all, which on 28/9 copied this guarantee among the conditions for a future adoption with exactly the right rule — «we verify it in a smoke test instead of assuming it, because a promise is a promise until you measure it» — and for whom the content of a report would be art. 9 GDPR data in a zero-knowledge product; then every product that sends the front texts which must leave no trace (nis2 voce, core/081, adp, dfm) and anyone who has to sign a DPA as an art. 28 sub-processor. How to adopt it: nothing in the clients, it is a guarantee that holds by itself; to re-run the measurement on your own installation launch fronte/prova_ritenzione.py (standard library only, no GPU, 4.1 s), with --mutazioni to watch it bite.

29/9/2026 — The front doesn't change the audio: parity with the engine measured byte for byte delivered, not online

The front sits between the products and the engine to add key, quota, log, honest queue and the OpenAI contract. Everything it adds is garnish: the audio that comes out must be the engine's, and until today nobody tested it. The difference between the two arms of the bench — the bare engine and the front — could therefore be read in two opposite ways: «the front costs N milliseconds» or «from the front the voice sounds different», with no number saying which of the two. Measured on 29/9: for the same request the front delivers the same bytes as the engine, on block (POST /tts, formats pcm16k and wav) and on stream (POST /tts-stream, pcm16k), with the same seconds of audio. The front's overhead is therefore pure overhead and it is measured: p50 10.7 ms on block and 3.2 ms on stream, over 6 pairs at one request at a time. The only admitted difference is long text, and the front declares it: beyond the engine's ceiling it splits at sentence end and puts 600 ms of silence between the pieces — tested byte for byte (audio = piece + silence + piece, the pieces asked of the bare engine; exactly 19,200 bytes of silence in pcm16k, not one more) and always accompanied by X-AgileTTS-Pezzi. A declared difference is a contract; the same difference undeclared is the audio changed on the sly. Tested by banco/tts/prova_parita_fronte_motore.py (36/36, four mutations that bite the right section and no other).

Who it is for: every product moving from the engine to the front with the additive patch — core/081 and adp first of all, then nis2 voce, vocea, dfm — because now the cost of the move is a number of milliseconds and not a doubt about the voice; and the owner, because the blind listening is only worth something if what comes out of the front is the same voice the bench measures on the engine. How to adopt it: nothing in the clients, it is a guarantee that holds by itself; whoever wants to redo the measure on their own installation runs banco/tts/prova_parita_fronte_motore.py (standard library only, no GPU), which prints the p50 and p95 overhead in milliseconds for block and for stream.

29/9/2026 — A stream that dies halfway is no longer invisible to the service delivered, not online

In streaming the 200 headers go out before the audio: from that moment on a failure can no longer become a 5xx. Measured on 29/9 with an engine that stalls after the headers: the front let the exception propagate and the handler wrote an HTTP/1.1 500 inside the already-open chunked body — noise in the audio of whoever was listening — while the log kept a http: 500 INTERNAL row with no audio_ms and nothing saying the client had been left halfway through the sentence. Now the front keeps the 200, no longer writes a second response inside the stream, leaves the chunked body without the final chunk (the only honest signal left after the headers) and logs troncato: <exception type>, never the text. The request counts as an error in the usage metering and produces no rtf samples: measuring the engine on half a sentence would be lying. pagina_uso.py has the "stream troncati" column and a threshold-free alarm: one is enough. Not truncations — and they close cleanly — the client that walks away (client_disconnect), the failed piece (pezzo_fallito) and the engine's anti-runaway ceiling. Tested by fronte/prova_stream_troncato.py (19/19, three mutations).

Who it is for: whoever listens to AgileTTS while it speaks rather than to a finished file — core/081 and adp/adp_brain first of all (a sentence cut in half on the phone or on the avatar is heard immediately), then nis2 voce, vocea, dfm — and the regia, which sees how many sentences died halfway and on which exception. How to adopt it: nothing in the clients, but it is worth putting in your own code: a chunked body that does not close with the final chunk is a truncated stream and must be treated as an error, not as a finished sentence (IncompleteRead with urllib/requests; by hand, the final 0\r\n\r\n).

28/9/2026 — The front measures the engine's speed and says whether the honest-queue ceiling is still honest delivered, not online

The honest queue refuses with 503 BUSY + Retry-After beyond max_coda in-flight requests (3). Measuring it on 28/9 showed that this ceiling is a fixed number in front of an engine whose speed changes: it is honest only while max_coda × rtf < 1, and with the engine measured on 21/9 (rtf 1.41 already at one synthesis at a time) the front would accept three conversations and serve all three worse than real time. The front now measures the engine rtf on every request it serves (from the meta's synth_s, or from wall time only if the request was alone: under load it would measure the others' wait) and says it in GET /health → coda_onesta {max_coda, in_volo, rtf_motore {p50, p95, campioni, minimo_campioni, da_synth, da_parete}, prodotto, onesto, regola}, with onesto empty until there are enough samples. The log carries the sample (rtf_motore, rtf_fonte) and pagina_uso.py raises the alarm when rtf p95 ≥ 1 or max_coda × rtf p95 ≥ 1. The ceiling does not move by itself: changing it is a decision for the regia. Tested by banco/tts/prova_tetto_onesto.py (36/36).

Who it is for: whoever sends conversations rather than files — core/081 and adp first of all (a phone call served at rtf > 1 accumulates delay at every turn), then nis2 voce, vocea, dfm — and the regia, which sees the ceiling turn dishonest before a user hears it. How to adopt it: nothing in the clients; whoever watches reads coda_onesta.onesto from /health or runs pagina_uso.py --json uso.json --riga --health …, which exits 2 if there is an alarm.

28/9/2026 — The per-request log does not carry the text even when an internal error quotes it delivered, not online

On 500 INTERNAL the per-request log kept err = the representation of the exception, repr(e). Many library exceptions carry in their message the string they tripped on (UnicodeEncodeError is the classic example): along that service path the client's text could end up on disk. Measured on 28/9 with an exception quoting the text: the line contained the whole sentence. Now err carries only the type ("ValueError", "UnicodeEncodeError", …); the full representation stays in the 500 response to the caller, who already has their own text. Tested by fronte/prova_registro.py.

Who it is for: every product sending texts that must leave no trace — vocea first of all (zero-knowledge, GDPR art. 9 data), then nis2 voce, core/081, adp, dfm. How to adopt it: nothing to do; the field is documented in docs/REGISTRO.md and the contract line in docs/API.md §3.

27/9/2026 — Lexicon extension: dates spelled out, decades and centuries, dotted acronyms, fractions and negative numbers available client-side

normalizza_it.py now also reads dates spelled out («il 25 dicembre», «25 dicembre 2026»), decades and centuries in figures («gli anni '90» → «gli anni novanta», «nel '68» → «nel sessantotto», «il 1400» → «il quattrocento»), dotted acronyms spelled out («U.S.A.» → «u esse a», «R.S.A.» → «erre esse a», «D.O.C.» → «di o ci»), fractions with one digit per side («1/2» → «un mezzo», «2/3» → «due terzi», «3/4» → «tre quarti») and the minus sign before a number («-5 °C» → «meno cinque gradi»). The cases in casi.json grow to 241 (217 → 241, L218–L241; this turn 231 → 241, L232–L241), verifica_lessico.py 241/241.

Who it is for: core/081 and adp (they read dates, decades, centuries and acronyms without guessing), nis2 voce, vocea, dfm. How to adopt it: as for the normalizer — copy normalizza_it.py and call normalizza(testo), or on the front AGILETTS_FRONTE_NORMALIZZA=1.

27/9/2026 — Command-line client for the front (obj. 6) delivered, not live

client/agiletts-cli.py: a command-line client in pure standard library (argparse/json/urllib) that wraps client/agiletts_client.py; speech for block or streaming synthesis to file (-o, with --voice/--model/--format/--speed/--stream), models/model/voices/usage to read models, voices and usage; the key is given with --api-key or AGILETTS_API_KEY (or AGILETTS_KEY), --base-url points to the front; the front's errors go to stderr with a non-zero exit. Tested 11/11.

Who it is for: whoever wants to try or automate synthesis without writing Python code (regia, testers, vigilance scripts) and every product that integrates AgileTTS from a script or a pipeline. How to adopt it: nothing until the front is online; then python3 client/agiletts-cli.py speech "testo" --voice serena -o saluto.wav (key with --api-key or AGILETTS_API_KEY).

27/9/2026 — Adoption guide for products delivered, not live

docs/ADOZIONE.md (and docs/ADOZIONE.en.md): the step-by-step guide to move a product onto the OpenAI-compatible front one at a time, with the owner's "go" and the rollback ready — URL (https://api.agiletts.agile.software/v1), key (Authorization: Bearer <key> or X-Api-Key), format (response_format/format), timeout (8 s nexus reference, ≥ 10 s until it's WARM), 503 BUSY + Retry-After and declared fallback (never silent), "go"/rollback sequence and the table of the 5 products (nexus-voice-ms, adp_brain, nis2 voce, vocea, dfm).

Who it is for: the integrators and customers who will have to move from ElevenLabs (or the engine) to the front. How to adopt it: it's documentation, read docs/ADOZIONE.md (or docs/ADOZIONE.en.md); at injection each product follows section 6 ("go" and rollback) one at a time.

27/9/2026 — Voice catalog with tracked consent and refs encrypted at rest (obj. 7) delivered, not live

The catalog becomes a «consented catalog»: provenienza.schema.json (type persona/sintetica/blend, consent with protocol/date/expiry/uses/revocation, scope, reference with sha256/duration/encrypted, fallback policy); audit_catalogo.py raises violations for a person without consent, expired or revoked consent, scope outside the allowed uses, a person in the public catalog, an unencrypted ref, a different sha256, audio tracked in git; cifra_ref.py encrypts the refs at rest (AES-256 + HMAC, key on a file, never printed). Tested 16/16.

Who it is for: the regia and whoever keeps the catalog on the GEX44. How to adopt it: nothing for clients; on the engine audit_catalogo.py --voices … --radice … --repo … and the refs are encrypted with cifra_ref.py cifra ref.wav chiave.file.

27/9/2026 — Documentation and site in Italian and English (bilingual delivery) delivered, not live

The front's contract and help in two languages: docs/API.en.md, docs/NOVITA.en.md, client/ESEMPI.en.md and the site pages in English with the it/en selector; prova_docs.py and prova_sito.py check the consistency. Proper names, brands and paths stay identical in English.

Who it is for: integrators and clients working in English (nis2 voce, vocea, dfm). How to adopt it: nothing; one opens the page with the EN selector. The Italian sources stay the truth.

26/9/2026 — Example client in the standard library (obj. 6) delivered, not live

client/agiletts_client.py: a Python client in pure standard library (block and streaming speech, models, voices, usage; the front's errors become Errore exceptions with http/code/message/retry_after/extra); client/ESEMPI.md with Python, curl and fetch examples. Tested 27/27.

Who it is for: nis2 voce, vocea, dfm and every Python product without the OpenAI SDK. How to adopt it: nothing until the front is online; then from agiletts_client import AgileTTS or the examples from ESEMPI.md.

22/9/2026 — Usage page in json with alarms and the status line (front end v0.5, ob. 10) delivered, not live

pagina_uso.py --json uso.json --riga writes the same measures of the page in json with the allarmi list and prints a status line ready for dico serverino "AGILETTS: …"; it exits 2 if there is at least one alarm. Absolute alarms: external fallbacks (never silent), 503 BUSY, failed pieces, 403 KEY_ROTATED per key, grace ended with requests still on the old key, a key that stopped calling, a mute front-end /health. Threshold alarms: errors ≥ ×2 and first-byte p95 ≥ ×1.5 compared to the hours before (with at least 10 requests per window).

Who it is for: regia and vigilance. How to adopt it: nothing for clients; it is the box's evening line.

22/9/2026 — Usage page: comparison with the previous hours (front end v0.5, ob. 10) delivered, not live

The «Compared to the previous N hours» section shows the same window shifted back with requests, ok, errors, first-byte and total p50/p95, audio, fallbacks, regenerations and 503 BUSY, with the Δ colored (red worsens, green improves). Below, new keys and vanished keys; in the per-key table, Δ requests and Δ errors.

Who it is for: regia and vigilance. How to adopt it: nothing; it is the page generated by the box.

22/9/2026 — Usage page with key rotations and split texts (front end v0.5, ob. 6 and 10) delivered, not live

«Key rotations» section: keys in grace (expiry and new key), grace ended, requests on the old key and 403 KEY_ROTATED per key; clear status and «used this month (family)» for each key. The log now records the rejections of a recognized key with the id, even on GET.

Who it is for: whoever manages keys during a rotation. How to adopt it: nothing for clients.

22/9/2026 — Long texts up to 4096 characters, split at sentence boundaries and rejoined (front end v0.5, ob. 6) delivered, not live

The front end accepts up to 4096 characters (over that, 413); above 2000 it splits the text at sentence boundaries, asks the pieces to the engine one after another and rejoins them with 600 ms of silence. It applies in block and streaming; response with X-AgileTTS-Pezzi and merged meta. A piece in error = error of the whole request. The engine does not change.

Who it is for: whoever sends long texts (nis2 voce, vocea, dfm). How to adopt it: send the whole text; under 2000 characters nothing changes.

21/9/2026 — Key rotation with a grace period (front end v0.4, ob. 6) delivered, not live

rotate <key_id> [grazia_ore=24] issues the new key and puts the old one in grace: it keeps working, but every response carries X-AgileTTS-Chiave-Scade and X-AgileTTS-Chiave-Nuova; once the grace ends it answers 403 KEY_ROTATED with rotated_to. The monthly quota follows the family (rotating does not gift characters).

Who it is for: regia and front-end clients — the key is changed without downtime. How to adopt it: at rotation, move the secret into the product's box within the grace.

21/9/2026 — Per-voice fallback policy, declared by the catalog (front end v0.3, ob. 8) delivered, not live

Each voice can have in provenienza.json a policy sovrana (default) or dichiarato with verso. On every 5xx error the front end answers with error.ripiego and X-AgileTTS-Ripiego and writes ripiego in the log. The front end never calls an external provider: the decision is the owner's per voice, the execution stays with the client.

Who it is for: core/081 and adp (today they fall back silently), owner. How to adopt it: on 5xx read X-AgileTTS-Ripiego and fall back only if consentito.

21/9/2026 — Front end v0.3: OpenAI contract tested without the SDK (ob. 6) delivered, not live

A client written for OpenAI works without changes: model: tts-1 | tts-1-hd | gpt-4o-mini-tts are aliases of agiletts; without response_format with an alias it returns mp3; GET /v1/models/{id}; new formats opus, flac, aac made by the front end with ffmpeg; stream_format: sse → honest 400.

Who it is for: nis2 voce, vocea, dfm. How to adopt it: at the «go» they change base_url and key, not the code.

21/9/2026 — Usage page with the «engine» section (ob. 10) delivered, not live

pagina_uso.py also reads from the log queue wait, 503 BUSY, regenerations, short phrases, p50 similarity, capped stream; per-key table of the last hours. Never texts nor keys in clear.

Who it is for: regia, then each product for its own key. How to adopt it: pagina_uso.py --health … --out uso.html and publish from the serverino.

21/9/2026 — Front end v0.2 aligned with the engine's news delivered, not live

/v1/models says that the 0.6B runs in production; /health.motore reports coda, stream_tetto, guardia_identita; on block the full meta passes, on stream the queue and cap headers; an engine 503 busy becomes 503 BUSY with the engine's Retry-After. Normalizer switch AGILETTS_FRONTE_NORMALIZZA=1 (default 0).

Who it is for: everyone at injection and the regia. How to adopt it: nothing today.

21/9/2026 — Italian text normalizer available client-side

normalizza_it.py (standard library only, one normalizza(testo) function) turns into words what the model today guesses: POD/PDR/IBAN/tax-code codes, amounts, percentages, dates, times, phones, units, ordinals, roman numerals, acronyms, accents on homographs. Measure: 200/200 on the cases of banco/lessico/casi.json.

Who it is for: core/081 and adp right away (they read POD, IBAN, amounts, dates). How to adopt it: copy normalizza_it.py and call normalizza(testo) before synthesis; it will enter the engine after the A/B with ears.

21/9/2026 — Anti-runaway cap on the stream delivered, not applied

Every stream has a cap = min(max_s of the voice, max(3 s, 1.0 s × words + 3 s)): beyond it, the engine closes the stream, writes STREAM_RUNAWAY and counts in /health.stream_tetto; the X-AgileTTS-StreamTetto header says the cap in advance. Limit: a stream closed by the cap is truncated audio, not an HTTP error.

Who it is for: core/081 and adp (no 18 s mumbles for 4 words), regia. How to adopt it: nothing client-side.

21/9/2026 — Honest engine queue delivered, not applied

Each request measures its wait (queue_wait_s, queue_ahead, X-AgileTTS-Queue header, coda block in /health). Only if the regia turns on AGILE_TTS_MAX_CODA=N (default 0), beyond N requests the new one receives immediately 503 busy + Retry-After instead of waiting silently.

Who it is for: core/081 and adp (knowing if the slow first byte is the GPU or the line), regia. How to adopt it: nothing while MAX_CODA stays 0.

21/9/2026 — Identity guard with a minimum duration delivered, not applied

On /tts the guard always measures but does not regenerate on short phrases (fewer than 3 words or audio under 1.2 s), where the meter does not distinguish the person; from 3 words up everything identical. New identity_short field in the meta and guardia_identita block in /health. Patrizia remains a case apart (fails even on long phrases).

Who it is for: core/081 (the «Sì.»/«Pronto?» drop from ~3 s to ~0.7-1 s), adp. How to adopt it: nothing client-side.

18/9/2026 — «always warm» patch extended: the /tts path is warmed too delivered, not applied

The mute synthesis after each load goes through both paths (stream with CUDA graph and block with eager decode, which do not share the compile). Periodic rewarm off by default (AGILETTS_PRECALDO_OGNI_S) to measure the «cold after an hour». /health.caldo gains riscaldo_vie, riscaldi_periodici, ultimo_riscaldo.

Who it is for: core/081 (first audio after a restart), adp. How to adopt it: nothing client-side.

15/9/2026 — OpenAI-compatible front end v0.1 delivered, not live

POST /v1/audio/speech, /v1/models, /v1/voices, /v1/usage with ats_ keys per product and tenant, scope, quota, usage metering, log without texts, X-AgileTTS-Via, honest queue 503 + Retry-After; historical pass-through /tts and /tts-stream with key.

Who it is for: all future clients and today's at injection. How to adopt it: key from the box and base_url with the OpenAI client.

15/9/2026 — «always warm» patch delivered, not applied

After each load the engine builds the voice prompts and does a mute synthesis: the first /tts-stream at 8.3 s after a restart disappears, and the 2 s at the first request of each voice.

Who it is for: adp (avatar), core/081. How to adopt it: nothing client-side; the restart lasts 60-90 s more.

14/9/2026 evening — The 0.6B model runs in production (regia, reversible) live

AGILETTS_MODEL → Qwen3-TTS 0.6B: 2 GB less on the shared board; the identity guard thresholds stay tuned on the 1.7B, so expect more regenerations until the identity banco retunes them.

Who it is for: all of today's clients. How to adopt it: nothing; whoever compares the voice should know the model changed.

14/9/2026 21:15 — Attack gate live

Removed the semi-vocalic breath that «serena» emitted at the head of every phrase, by structure and text, on /tts and /tts-stream. New attacco field in X-AgileTTS-Meta (taglio_ms, scarti, motivo), trim_attacco in the body to turn it off and in /health.

Who it is for: core/081 and adp. How to adopt it: core must make its downstream _agiletts_trim_head inert (now it cuts twice); if a phrase loses its first syllable, report text and voice.

14/9/2026 14:35 — Resident model (regia) live

AGILETTS_WARM=1, AGILETTS_IDLE_UNLOAD_S=0: no unload after 180 s; first byte 333 ms instead of 1655 ms after the unload.

Who it is for: everyone (especially core/081: the first 60-70 s audios disappear). How to adopt it: nothing; the 8 s timeouts on the core side can stay.