On September 10, 2026, OpenAI released GPT‑Live‑1 into its API — a model that can listen and speak at the same time. Listen without going silent. Be interrupted — and keep going. Ask again. Circle back. Never wait for its turn. At first glance this is another improvement to voice assistants. In fact, it is a shift in the very structure of conversation between a human and a machine.
Until now, voice AI worked like a telegraph: a line, a pause, an answer, a pause, the next line. GPT‑Live‑1 moves the mechanics of real conversation — interruptions, overlaps, half-finished sentences, small acknowledgements — inside the model itself. And that changes not the quality of the answer, but the nature of the interaction.
AI no longer has to wait its turn
Until recently, almost every voice system was built the same way: speech is recognized, converted to text, processed by a language model, converted back into speech. Technically it worked. But the logic stayed sequential: the system had to wait for the human to finish speaking before it began to reply.
Humans do not talk like that. We interrupt. We pause to think. We start a sentence and abandon it halfway. We listen and formulate a response at the same time. We react to tone, not only to words. And for decades, these subtleties remained outside the reach of voice interfaces.
GPT‑Live‑1 brings that mechanics closer to the model. OpenAI describes it as full‑duplex — listening and speaking simultaneously, handling interruptions, sustaining continuous conversation. More complex reasoning and tool calls can be delegated to other models.
This is no longer an improvement to Text‑to‑Speech. It is a different interface.
From a voice interface to continuous interaction
At first glance, the market looks like it is simply moving toward more realistic voice assistants. But something more fundamental is happening.
We are moving from a model of “AI that answers requests” to a model of “AI that is continuously interacting with a person and with the environment.”
In the first, the human initiates action: asked — got an answer — asked the next question. In the second, the human and AI are in continuous interaction. AI can listen, understand whether the person has finished or is still speaking, react to interruption, ask clarifying questions, use tools, receive new data, perform actions, return with results, and continue the conversation in the context of the previous steps.
The interface starts to resemble not a chat, but an operator.
And it is no longer only OpenAI
GPT‑Live‑1 is not developing in a vacuum. Google is building the Gemini Live API — a realtime, multimodal layer that ingests continuous streams of audio, video, and text, and supports spoken responses, interruption, affective dialogue, tool use, and proactive audio. ElevenLabs is building a stack that long ago moved past ordinary Text‑to‑Speech: Eleven v3, realtime speech‑to‑text, realtime conversational speech, voice agents, telephony, knowledge, tools, and SDKs/APIs for building agents. Cartesia is taking an even more infrastructure‑oriented route, combining realtime TTS, streaming speech‑to‑text, turn boundary detection, and voice agents.
A distinct technological layer is forming — AI interaction infrastructure.
From Text-to-Speech to AI that can hold a conversation
This is an important transition. The first generation of voice technology answered: how do we turn text into human speech? The next: how do we recognize human speech? Then: how do we make speech emotional and natural? Now the question is different: how do we make AI a full participant in a conversation?
ElevenLabs is especially telling here. The company started with exceptionally strong Text‑to‑Speech, but gradually expanded its stack to Speech‑to‑Text → reasoning → conversation → voice → agents → tools → telephony. Its realtime speech‑to‑text, Scribe v2 Realtime, is built for live interaction and claims sub‑150 ms latency. That is an important signal to the market: voice is becoming not a media format, but an interface for agents.
But here is where it gets interesting
We can teach AI to hear, speak, see, recognize emotions, detect when a person has finished a sentence, react to interruption, use tools. But the next question arises: what exactly does AI understand?
Suppose a leader says: “Move tomorrow's appointment and tell the doctor.” A modern realtime AI can already hear that sentence perfectly. It can even understand its general meaning. But to perform the action, it needs to know far more. Which appointment exactly? Which doctor? Which branch? Which rules apply when rescheduling? Is this user allowed to do it? What resources will be freed? Who else needs to be notified? Which KPIs or obligations are affected? Which system is the source of truth? What must be written to the audit trail?
This is where the boundary lies. Speech recognition is not yet enterprise understanding.
AI can hear you and still not understand the company
This is, we believe, one of the central questions of the next phase of AI. A model can have a huge context window. It can have access to millions of documents. It can have excellent reasoning. It can hold a conversation with almost no latency. But that still does not mean it knows what “customer” means in this particular company; how “order” differs from “request”; who owns the process; which rules apply; which actions a specific person is allowed to take; how a particular KPI is calculated; which system holds the authoritative state.
That is why enterprise context matters more than the interface itself.
From AI that knows the data to AI that understands the enterprise
This direction is also visible in the evolution of modern data and AI platforms. A unified data store, descriptions of documents, terms, KPIs and methodologies, and an AI agent working directly with that layer — that is already far beyond an ordinary chatbot.
But the next level lies deeper. There is a big difference between “AI understands the company's data” and “AI understands the company.” In the second case, we are not talking only about data. We are talking about people, objects, processes, roles, responsibility, rules, knowledge, systems, events, decisions, and the current state of the enterprise. This is where Enterprise Ontology appears.
The interface is only the top layer
If you look at the emerging architecture of AI systems, it increasingly resembles a layered structure. The interface is only one of its layers. Behind it stand interaction, orchestration, enterprise context, reasoning models, tools, the operational environment, and the evidence loop.
AI interaction layers — from perception to operational action.
And it becomes obvious: GPT‑Live‑1 is not an enterprise operating system. ElevenLabs is not an enterprise operating system. Gemini Live is not an enterprise operating system. They solve different parts of a larger problem. And that is good news.
Models are becoming replaceable
One of the most interesting architectural trends is that the interaction layer and the reasoning layer are beginning to separate. OpenAI explicitly describes GPT‑Live‑1 as a realtime interaction layer that can pass more complex reasoning and tool calls to other models.
This means the architecture might look like this: Voice → GPT‑Live → Enterprise Orchestrator → Enterprise Model → GPT / Claude / Gemini / local model → tools / workflows → enterprise systems.
And this is a very important principle. Intelligence should be replaceable. Enterprise context should not.
That is why the formula “own the enterprise model, own the orchestration, keep the intelligence replaceable” becomes even more relevant.
What is actually becoming the new operating layer
Looking at the market today, we can see several rapidly developing layers.
- Perception. Speech, vision, video, sensors.
- Interaction. Realtime voice, full‑duplex conversation, multimodal dialogue.
- Reasoning. Large language and multimodal models.
- Context. Knowledge, data, the enterprise model.
- Orchestration. Agents, tools, permissions, workflows.
- Action. Applications, APIs, business processes.
- Evidence. Audit, outcomes, runtime state.
- Evolution. Updating the enterprise model based on what happens in reality.
Today the market is rapidly developing the first three or four layers. But enterprise context, orchestration, and operational state will determine whether AI can become a real participant in the business.
So voice AI is not the final product
There is a fundamentally important conclusion here for us. We should not turn Enterprise AI into a “voice assistant for business.” That is too narrow. Voice is only one interface. Tomorrow the interface might be text, voice, a camera, a screen, a wearable, AR glasses, a car, a robot, an embedded interface of an application, or a fully autonomous agent.
We do not need to own the interface. We need to own the context within which any interface can operate.
From conversation to action
And here GPT‑Live‑1 fits our concept especially well. Today AI talks to a human. The next step is AI using tools. The next is AI executing business processes. And after that — AI becoming an operational participant in the enterprise.
But that requires an enterprise model. Because action without context is just automation. And action inside a model, processes, permissions, and rules is operational activity.
Enterprise Design becomes the connecting layer
That is why our architecture Understand → Design → Operate → Augment gains another very practical meaning. Understand how the enterprise actually works. Turn that into an Enterprise Blueprint. Turn the Blueprint into a digital operating environment. Let AI work inside that environment. And then GPT‑Live‑1 becomes not a competitor to that concept, but an example of a new interaction layer.
What changes for the enterprise
Imagine a leader who says in the morning: “Show me where we lost revenue yesterday.” AI does not just answer with voice. It understands what “revenue” means in this company; understands the structure of the business; knows the KPIs and their methodology; identifies the relevant events; analyses the data; explains the causes; suggests actions; and, if authorized, launches the corresponding workflow; and records the outcome.
Then the leader interrupts: “Stop. Compare it with last week too.” AI does not need to “restart.” It continues the conversation. Because this is no longer a series of requests. This is continuous interaction with the operating system of the enterprise.
The most important shift
So we would frame it this way. First generation: AI can answer. Second: AI can converse. Third: AI can act. Next: AI can act inside the enterprise. And the last one requires not just a smarter model. It requires a model of the enterprise itself.
From “talking AI” to Enterprise AI
GPT‑Live‑1 matters not because AI now sounds almost human. That is already becoming an expected characteristic. What matters far more is that the boundary between interface and action is getting thinner. AI is starting to be perceived not as an application you open and ask a question to. It is becoming a permanent layer of interaction.
And the more natural this layer becomes, the more important the question becomes: inside what world does it operate? For a personal assistant, that world may be the user's life. For enterprise AI, that world must be the enterprise itself.
Our conclusion
Full‑duplex makes AI a natural interlocutor. Enterprise Design makes AI a participant in the work of the enterprise. The first solves the problem of interaction. The second solves the problem of meaning, context, permissions, and action.
That is why the next stage of corporate AI is not simply AI that can speak. And not even AI that can act. It is AI that understands where exactly it acts. In the enterprise. With its objects. Processes. Rules. Responsibility. Knowledge. Data. Systems. And current state.
That is why we believe the future of enterprise AI is not built around a single model. It is built around a model of the enterprise itself.
Understand → Design → Operate → Augment. The future of AI is not better conversations. It is AI that can operate inside the enterprise.

