SeedRealtime icon
SeedRealtime
#84 in Developer Tools
4.4/5
« An AI model that sees, hears, and speaks simultaneously within a single architecture, enabling real-time video conversations with responsiveness close to that of a natural human exchange »
Free 6924

SeedRealtime: ByteDance keeps the mic and camera open, and the AI talks while it watches

SeedRealtime is ByteDance's full-duplex audio-visual AI model: it watches through the camera, listens and speaks at the same time, instead of waiting for its turn like a walkie-talkie. The only place to try it is Doubao, the company's consumer assistant, where voice and video conversation costs nothing. No API and no price list have been announced for developers. Everything else, from parameter count to benchmarks, remains behind closed doors.

Pros
  • Watches, listens and speaks at once, no turn-taking
  • Speaks up on its own when the scene changes
  • Tracks several speakers in the same room
  • Free to try inside the Doubao app
Cons
  • Doubao-only for now, with no developer API
  • Signing up usually calls for a Chinese phone number
  • Still young: no technical report or detailed benchmarks yet

One brain instead of three chained modules

SeedRealtime fuses audio, video and text in a single end-to-end architecture, where most voice-first AI assistants chain speech recognition, a language model and text-to-speech, adding latency and losing context at every handoff. Even the decision to speak or stay quiet happens inside the model, with no external voice-activity detector.

The payoff shows in conversational rhythm. In ByteDance's own human evaluation, pacing problems, cut-off replies, slow answers after a pause, false triggers from nearby speech, were cut in half against cascaded systems. It is the full-duplex) principle applied to AI: both sides can talk at once, like on a phone call.

SeedRealtime in the wild: noisy dinners, museums, coffee machines

ByteDance published several scenarios where the model works from a continuous audio-video stream. You flip through a document in front of the camera and it stops you at the right chapter. The most striking demo has it binding each voice to a face around a table.

SituationWhat SeedRealtime does
Noisy group dinnerMatches names, faces and voices, credits each opinion to the right person
Museum visitPipes up on its own when the artifact you asked about appears
Unfamiliar coffee machineTalks you through the steps and corrects based on what it sees
Restaurant abroadHelps decode the menu and the waiter's explanations live
Reading a paperSpots the target page as you flip through

Doubao, the only door into SeedRealtime

The model is fully rolled out in Doubao, through the app's call feature, in voice and video, at no charge. The catch: registration usually requires a mainland China mobile number (+86), a real hurdle outside the country, and the app targets the Chinese market first.

Developers will have to wait. There is no Volcano Engine or BytePlus endpoint, no open weights and no technical report (ByteDance is keeping its cards close for now). The stated roadmap covers lower latency, sharper speaker tracking and tool-connected tasks such as lookups and bookings.

Frequently asked questions

Is SeedRealtime free?

Yes, SeedRealtime costs nothing inside ByteDance's Doubao app, in both voice and video conversation. No paid tier, published quota or billed API has been announced. Access runs exclusively through the consumer app, so a valid Doubao account is the only requirement.

Can you use SeedRealtime outside China?

With difficulty for now. Doubao registration usually asks for a mainland China mobile number (+86), and the app is built for the Chinese market. From elsewhere, the model is mostly visible through the demo scenarios published by ByteDance's Seed team.

Does SeedRealtime have a developer API?

No. ByteDance has released no Volcano Engine or BytePlus endpoint, no open weights and no technical report for this model, so third-party teams cannot integrate it today. The roadmap mentions further work, but gives no timeline for opening access to developers.

How is SeedRealtime different from ChatGPT's voice mode?

Native video is the dividing line. SeedRealtime fuses audio, vision and text in one full-duplex model that can comment on what it sees while you are still talking, and even take the initiative. OpenAI and Google push similar real-time capabilities, but ByteDance claims this end-to-end audio-visual fusion as its signature.

Verdict: Half as many cut-offs and false starts, if ByteDance's numbers hold: anyone building or studying camera-and-mic assistants will read SeedRealtime as a signpost for where machine conversation is heading, and Doubao regulars can already chat with it for free.

★ Featured AI Tools ★
AI Alternatives for SeedRealtime