Nvidia has introduced Nemotron 3 Diarization, a new AI model capable of identifying up to eight speakers in real time. This 100 million-parameter model can detect which speaker is talking at any moment during a conversation, enhancing transcription and communication tools.
According to The Decoder, Nemotron 3 Diarization is available for free, making advanced speaker diarization technology more accessible to developers and businesses. The model’s ability to handle multiple speakers simultaneously marks a significant step in AI-driven audio processing.
For Japanese markets, where multilingual communication and precise transcription play a growing role in financial and tech sectors, Nvidia’s release could boost innovations in AI-powered communication tools and services.
