OpenAI Testing AI-Generated Voice Mimicry in Limited Private Preview

OpenAI is testing a new AI-based voice technology in an effort to explore its capabilities while keeping it out of the hands of potential bad actors.

Voice Engine "uses text input and a single 15-second audio sample to generate natural-sounding speech that closely resembles the original speaker," OpenAI said in a blog post last week. It can create new audio from text using a single person's voice for reference, translate existing audio into another language while retaining the original speaker's tone and accent, or create new audio in a language that's different from the original speaker's.

Voice Engine has been around for a few years; its technology underlies OpenAI's text-to-speech API and ChatGPT's voice querying capability that was introduced last fall. At the end of 2023, however, OpenAI decided to start priming Voice Engine for eventual public consumption, beginning with just a "small group of trusted partners."

OpenAI did not indicate how long this limited private preview will last, or when it expects to make Voice Engine generally available. It is purposely taking a slow and measured approach "due to the potential for synthetic voice misuse," it said. OpenAI took a similar tack when launching its Sora text-to-video capability in February, making it available to only select testers.

Nevertheless, the Voice Engine testing group has already begun applying the model to a few real-world applications across several industries. For instance, it's being used to help patients with speech conditions to communicate with their own voice, using old video recordings of them as the reference audio. Content creators are also using it to translate their assets into different languages, giving them a broader audience. Other examples, with real audio snippets of Voice Engine in action, are on OpenAI's blog post.

For all its capabilities, however, this technology, like Sora, is ripe for abuse. OpenAI said it is working to develop guardrails for Voice Engine to limit how much it can contribute to the spread of misinformation. For instance, Voice Engine prevents individual users from making their own voices from scratch. In addition, audio created by Voice Engine comes with "watermarks" to track each snippet's provenance, as well as how it's being used.

OpenAI said it gave its testing group access to Voice Engine under several stipulations to discourage them from abusing the technology. For instance, they are not allowed to use Voice Engine to impersonate another party without their explicit consent. Testers also have to disclose to their audience when they're using Voice Engine to create audio.

To avoid widespread harm caused by AI-generated audio in general, OpenAI makes several suggestions for policymakers and developers:

  • Enforce a "no-go voice list" to prevent the impersonation of well-known figures.
  • Avoid voice-based authentication for critical systems.
  • Develop ways for individuals to protect the ownership of their voice.
  • Spread broad public awareness of AI misuse.
  • Fast-track the development of technology that can trace the provenance of audio and visual media.

"We recognize that generating speech that resembles people's voices has serious risks .... We are engaging with U.S. and international partners from across government, media, entertainment, education, civil society, and beyond to ensure we are incorporating their feedback as we build," OpenAI said.

About the Author

Gladys Rama (@GladysRama3) is the editorial director of Converge360.

Featured

  • Digital cyberspace with particles and Digital data

    Report: AI Is Moving Faster than Data Trust

    AI agents are already in use or pilot at most organizations, but data visibility, governance and precision recovery capabilities have not kept pace, according to Veeam's new Data & AI Trust Gap report.

  • Glowing route lines merging into single gold pathway

    Microsoft Merges Copilot Apps Into a Single User Experience

    Microsoft recently announced it is consolidating its consumer Copilot app and Microsoft 365 Copilot app into a single destination, addressing a fragmented product strategy that has required users to navigate separate applications for AI chat and productivity tools.

  • CIS Sandbox Tutor Eli Blouin at Bentley University imagines a future technology learning partner that supports curiosity, critical thinking, feedback, and personalized learning.

    Student Voices: Technology as a Future Learning Partner

    A data analytics and marketing major at Bentley University and a distinguished lecturer at the university explore the future of "technology learning partners" — technologies that participate in the learning process by asking questions and giving feedback, not just information search results.

  • Silhouettes of business professionals stand against a blurred futuristic city skyline at night, with a glowing digital network data connection

    It's Time for Higher Ed to Get Serious About AI Strategy

    Without a coordinated strategy that involves multiple academic and administrative units across the entire campus, colleges risk wasting resources, duplicating efforts, and ultimately failing to deliver on the promise of deploying technology to improve learning and operations.