Summary
- Deepslate has raised €7.7 million in a seed round led by 42CAP, with Alstin Capital, SIVentures, and business angels also participating.
- Its models process spoken audio directly instead of routing every conversation through separate speech-to-text and text-to-speech systems.
- European hosting, self-hosting, and local-language performance form part of its pitch as voice AI moves into operational business systems.
Voice AI is becoming another test of whether Europe can build specialist models around its own languages, infrastructure, and deployment requirements, with Berlin-based Deepslate raising €7.7 million to expand a speech-to-speech platform designed around that market.
The seed round is led by 42CAP, with Alstin Capital, SIVentures, and business angels also participating. Deepslate plans to use the capital for model research, European training data, commercial expansion, and additional computing infrastructure.
Rather than relying entirely on the familiar sequence of converting speech into text, passing that text through a language model, and synthesising a spoken response, Deepslate trains systems that process audio more directly. Its founders argue that the approach can preserve information such as tone, emphasis, timing, and dialect that a transcript may flatten.
Current reporting puts the model’s response time at about 0.44 seconds in testing by Artificial Analysis, while Deepslate advertises support for 27 languages. The company also offers European hosting and deployment on customer-controlled infrastructure.
Voice systems have less room for delay
Latency is more visible in speech than in many other AI interfaces because people quickly notice an unnatural gap during a conversation. When a system pauses for too long, callers interrupt, repeat themselves, or assume the connection has failed, which makes milliseconds commercially relevant once voice AI moves into call handling and operational software.
The quality problem extends beyond speed because real conversations contain interruptions, incomplete sentences, hesitation, regional pronunciation, noise, and emotional cues. A text transcript captures the words but can lose some of the information carried by how they were spoken.
Direct audio processing is intended to reduce that loss, although it comes with its own technical demands around training data, inference cost, and evaluation. A model has to perform consistently across languages and accents rather than excelling only in the most heavily represented speech datasets.
European language coverage therefore forms part of Deepslate’s commercial pitch rather than only a localisation exercise. Large English-language datasets give global model providers an obvious advantage, while smaller languages, dialects, names, and locally specific terminology may require more deliberate collection and training.
The company says its technology is already being used in production with insurers, contact centres, and software platforms. Those deployments provide a tougher test than demonstration audio because operational systems have to cope with telephone quality, concurrent users, network conditions, business integrations, and failures that occur outside a controlled laboratory.
Hosting becomes part of the buying decision
Voice interactions can contain personal, commercially sensitive, financial, or health-related information, depending on the organisation using the system. Companies therefore need to understand where recordings and derived data travel, which suppliers can access them, and whether conversations are retained for training or monitoring.
Deepslate is using European hosting and self-hosting as part of its answer, although geography alone does not settle questions of sovereignty or control. Buyers still need to examine model ownership, software dependencies, cloud infrastructure, support access, and the jurisdictions applying to every part of the service.
Specialist providers also face a scale problem because training and serving speech models require substantial compute, while large model companies can spread those costs across broader product portfolios. Deepslate will have to show that narrower optimisation around voice, languages, latency, and deployment control creates enough advantage to offset that difference in resources.
The market does not necessarily require one model to dominate every speech application, however, because businesses may value a specialised system when local-language performance, infrastructure control, predictable latency, or integration matter more than access to the broadest possible general-purpose model.
As voice AI shifts from impressive demonstrations into contact centres, insurance workflows, software products, and other business processes, those operational characteristics become easier to measure. Customers can test whether the system understands callers, completes tasks correctly, stays responsive under load, and handles data according to their requirements.
Deepslate’s €7.7 million round is therefore financing a narrower proposition than building another general-purpose AI laboratory. The company is concentrating on the point where machines have to participate in spoken conversations quickly and reliably enough to become part of normal business infrastructure.












