NVIDIA open-sources Audio2Face, its AI facial animation model.

  • Audio2Face goes open source with SDK, v3.0 models, and training framework.
  • Official plugins for Autodesk Maya and Unreal Engine 5 make integration easy.
  • Includes Audio2Emotion and sample data for testing and customization.
  • Widespread industry adoption and a call for community collaboration on Discord.

AI facial animation technology

NVIDIA's decision to open-source Audio2Face marks a significant step for those creating digital characters with natural expressiveness. With this move, the company is encouraging more studios and developers to integrate AI-generated facial animation and lip-sync into video games, 3D applications, and immersive experiences without the usual access barriers.

The release includes the Audio2Face SDK , regression and diffusion models (version 3.0 ), and the training framework for fine-tuning behavior with custom data. The focus is on accelerating the adoption of AI-based avatars in sectors such as gaming, media, entertainment, and customer service.

What is Audio2Face and why does it matter?

Audio2Face transforms speech signals (phonemes, prosody, and emotional nuances) into curves and animation data that faithfully synchronize lips and expressions. This output can be used in real time or offline , covering everything from pre-recorded cinematics to live, dynamic interactions within a graphics engine.

For the player or viewer, the result is more believable expressiveness , with characters reacting consistently to the tone and rhythm of the audio, improving immersion in dialogue scenes, close-ups , and services with virtual assistants.

SDKs, templates, and tools available

The release includes the Audio2Face SDK , regression and diffusion models v3.0 , and the training environment needed to adapt the technology to different facial styles and rigs. Official plugins for Autodesk Maya (v2.0) and Unreal Engine 5 (v2.5) are also included , making integration into professional pipelines straightforward.

Audio2Face open source

In addition, complementary models such as Audio2Emotion , capable of inferring emotional states from audio, are available , along with sample datasets to start experimenting as soon as possible. For those seeking more information and resources, NVIDIA refers to ACE for Games , where the complete set of related tools is compiled.

Integration into 3D workflows

In existing productions, Maya and Unreal Engine 5 plugins make it easy to map Audio2Face output to facial rigs and combine it with handcrafted animation layers or capture systems. The SDK allows for process automation , building internal tools, and connecting AI with animation editors or rendering systems commonly used in studios.

The technology is optimized to run at high performance on modern GPUs (such as the RTX series), although the fact that the code is open source makes it easier to explore other deployment configurations and tailor adjustments to the needs of each project.

Complementary models and customization

With the released training framework , technical teams can refine models with their own phoneme tree, linguistic rules, and voice variety, or tailor the output to specific rig styles. The combination with Audio2Emotion opens the door to expressive nuances that better reflect the speaker's timing and intent.

For beginners, sample data allows you to validate audio channeling, test lip-sync, and evaluate the quality of transfer to rigs before investing in your own training corpus.

Industry Adoption

Audio2Face has already been integrated into tools and projects by studios and suppliers in the industry. Among the names mentioned are Codemasters, NetEase, Reallusion, Perfect World Games, GSC Games World, Convai, Inworld AI, Streamlabs, and UneeQ Digital Humans, indicating that the technology has matured in real-world environments.

  • reallusion incorporated Audio2Face in iClone y character creator, combining it with functions such as face puppeteering y AccuLip to fine-tune the lip-sync.
  • Survivals, Alien: Rogue Incursion Evolved Edition, optimized its pipeline facial animation to raise the immersion in virtual reality.
  • The 51 Farm applied it in Chernobylite 2: Exclusion Zone, reaching a level of realism superior to that of its first release.

Open community and collaboration

With the open-source code, NVIDIA invites developers, students, and researchers to contribute improvements, propose new features, and adapt the solution to specific use cases. The company also encourages participation in the Audio2Face Discord server , a meeting point for sharing progress and resolving technical questions.

The license change makes it easier for the community to experiment with heterogeneous workflows , from video games and VTuning to corporate virtual assistants, consolidating a codebase on which to iterate quickly and transparently.

With the opening of Audio2Face, the AI-guided facial animation ecosystem gains a significant boost: greater access , better integrations, and an adoption timeline that, judging by existing cases, has potential for both AAA productions and small teams seeking quality without starting from scratch.

15 RPGs with the most impressive customization options
Related article:
15 RPGs with the most impressive customization options

Add as preferred source in Google