June 16-24, OpenAI’s next-generation speech modelBidi-1(internal codename GPT-bidi-1) has experienced an intensive round of community leaks. Although OpenAI has not officially said a word yet, information from TestingCatalog, ChrissGPT and other channels has pieced together a fairly complete picture: a system that canListen and speak at the same time, support real-time translation, and provide three levels of intelligenceThe next generation of voice AI is coming soon.
What does "bi-directional" mean?
"Bidi" is the abbreviation of "Bi-directional", which is the voice mode of Bidi-1 and the current GPT-4oThe biggest architectural difference. currentChatGPTAlthough the voice mode supports interruption, it is essentially a "turn system" - after you finish speaking, it thinks and it answers. And true two-way audio means:The model can generate answers while listening to you.
This is a fundamental change in the user experience: just like chatting with a real person, you can respond with "hmm" in the middle of the other person's words, add in mid-sentence, and change your words while speaking - the model can understand and respond naturally, instead of waiting for you to "finish" before starting to process.
Three levels of intelligence: Instant/Medium/High
According to detailed information from TestingCatalog, Bidi-1 will be introduced in voice mode for the first time.Three optional intelligence levels:
- Instant: Extremely low latency, suitable for daily fast conversations
- Medium: Balance speed and depth
- High: Invoke stronger reasoning ability, suitable for complex problems
This layered design is similar to Anthropic's "Extended Thinking" switch - allowing users to choose between speed and quality. For developers, this means that voice capabilities at different costs can be invoked in different scenarios through APIs.
Real-time translation: API’s biggest selling point
On June 23, TestingCatalog further broke the news: Bidi-1 will have built-inreal-time translationcapabilities, and "this will unlock a large number of API use cases." The current Whisper+GPT+TTS translation pipeline requires hundreds of milliseconds of latency, while Bidi-1 promisesEnd-to-end speech translation in a single bidirectional audio stream——Without text relay.
If true, this will directly impact the existing simultaneous interpretation tools, real-time translation for video conferencing, and AI assistant markets for customer service call centers.
"Provisos" to note
- No official confirmation: As of June 26, OpenAI has not released any Bidi-1 model cards, pricing or release dates.
- The source of audio samples is questionable: The early test audio source shared by ChrissGPT is "Friends' Early Access" - the chain is opaque
- Language coverage unknown:What language pairs are supported by real-time translation? No public information
- European users may lag behind: TestingCatalog reminds EEA, UK, and Swiss users that they may gain access later.
Summary: The next inflection point of voice AI
The leaked information of Bidi-1 depicts a clear upgrade path: GPT-4o's turn-based speech will be replaced by true two-way dialogue, and real-time translation will become a standard API capability. Whether Bidi-1 is released next week or next quarter,The competitive bar for voice AI has been raised——Anthropic, Google and the open source community will all be forced to benchmark.
Want to learn about the latest developments in AI tools as soon as possible? access AI Dash Check out the latest AI tool reviews.
