Your own voice, from 15 seconds
A clean 15-second recording is enough: the announcement voice already on your switchboard, your best agent, or a studio take. Tone, tempo, feeling and speed are yours to set and to save as a preset bound to an agent. 40+ ready Turkish voices ship alongside it.
Turkish handled first, not translated into
Numbers and dates are read as words, so 1.250 TL is spoken and not spelled. The Turkish lowercase rule is applied properly — İ to i, I to ı — because the other way round the model reads nonsense. Brand and technical terms get a pronunciation dictionary, and pauses are placed at sentence ends rather than sprayed through the line.
It answers from what you gave it, and nothing else
Upload PDF, Word, Excel, HTML, CSV or plain text. Each customer sentence is matched to the passage that answers it, and the reply is grounded in that passage. With strict mode on, anything outside the knowledge base is declined politely — general knowledge, arithmetic and small talk included.
A turn in about a second and a half
Measured on real SIP calls: 0.1 of a second from the caller stopping to the transcript being ready, under 0.2 of a second to the model's first token, and about 1.2 seconds until the first audio is on the line. It was eight to nine seconds before the work that closed it.
Built for a telephone line, not a browser tab
The agent stops the moment the caller starts speaking. Room noise does not cancel the greeting. Audio is generated and streamed a sentence at a time, so even a long answer starts arriving immediately, and it speaks G.711 — the codec Turkish PSTN actually runs on.
Separation you can point at
Every customer is a workspace with its own voices, agents, knowledge base and SIP identity, separated at the application layer and at the SIP layer both. Usage is metered by the second against a ledger rather than a single editable number, and generation stops when the balance does.