Skip to content
TECH STARTUP & BUSINESS NEWS

I Got My Own AI Avatar Built by a $4 Billion Startup, and It’s Both Fascinating and Unsettling

When Alexandru Voica, head of corporate affairs at the video-generation startup Synthesia, sent a link this summer to the newest addition to the company’s public relations team, the result was immediately striking. It was an interactive virtual avatar modeled directly after him, trained specifically to answer common press inquiries regarding Synthesia’s core operations, functionality, and technology. Just a day prior, Voica had been participating on a panel where traditional PR professionals debated whether they were comfortable with pitches driven by AI-generated text. Yet, Voica’s digital twin went far beyond standard text generation, representing what felt like a definitive milestone in the application of artificial intelligence within modern public relations.

By September, Synthesia invited visitors to tour its newly established office space in New York. Originally founded in the United Kingdom, Synthesia stands as a prominent digital avatar enterprise, operating alongside competitors such as D-ID, HeyGen, and Colossyan. The company reached a remarkable $4 billion valuation earlier this year, building on previous milestones that included crossing $100 million in annual recurring revenue (ARR) alongside major backing from investors like Adobe.

Synthesia’s primary business model allows enterprise organizations to construct interactive training videos utilizing AI-driven avatars. Expanding upon this foundation, the company recently launched a product known as Roleplay Sessions. This platform enables employees to practice complex workplace scenarios—such as high-stakes sales pitches—with an interactive AI avatar that dynamically responds to their arguments and evaluates their performance in real-time.

During a visit to the new office opening, when representatives asked if there was any interest in creating a personal AI avatar, the response was immediate. The prospect of generating a digital twin seemed genuinely appealing, especially on a day when the outfit was well-chosen and personal presentation was entirely in order.

Prior to meeting this digital twin, personal sentiment toward avatars had largely been indifferent, though it was clear they would inevitably integrate into everyday online life. Observing social media platforms like Instagram, where creators frequently utilize digital likenesses to streamline content creation, highlighted a broader cultural shift. This underlying curiosity ultimately drove the decision to test the technology firsthand, marking the first time Synthesia created a digital avatar for a journalist or any external individual outside of company staff like Voica.

This specific interactive avatar was trained extensively on a previously published investigative story regarding why venture-backed startups experience higher rates of fraud compared to non-VC-backed companies. Consequently, the model is programmed to strictly answer questions pertaining to that specific piece of journalism, filtering out unrelated inquiries.

The creation process took place inside a compact film studio nestled within Synthesia’s office, where numerous photographs were captured alongside a two-minute voice recording. After providing formal consent, the digital version of the author was born. The engineering team produced multiple variations: a personal avatar designed to read verbatim scripts, complete with options for wearing or omitting glasses, and two interactive avatars capable of listening and talking back, likewise offered in both eyewear configurations.

Once the training material was selected, the engineering team deployed a complex technology stack to build the interactive agent. The system combines voice-to-text conversion, agentic language processing, natural language understanding, text-to-voice generation, and computer vision models. While the author’s avatar incorporates Synthesia’s proprietary video and voice models, the broader enterprise platform allows customers to select alternative technologies from specialized AI labs such as Cartesia, ElevenLabs, Google, or OpenAI. Furthermore, corporate clients maintain the flexibility to host their custom avatars on preferred cloud infrastructures or contract Synthesia for hosting services.

Operationally, the voice-to-text model transcribes spoken user input into readable data, the agentic language model interprets meaning and determines appropriate actions, the text-to-voice model converts the resulting response into natural-sounding audio, and finally, a specialized video model animates the avatar to synchronize facial movements with speech.

In total, Synthesia offers three distinct product categories: a traditional video-creation platform where users input scripts for static avatars to recite; an agentic platform designated for surveys and interactive roleplay sessions; and an API platform allowing developers to integrate Synthesia’s core video and voice models with third-party software to build custom interactive agents.

Constructing the custom avatars took the Synthesia team a matter of days. Initial testing began with the personal avatar by typing a standard, generic script to evaluate the accuracy of the synthesized voice. The test script focused on the arrival of autumn in New York, a favored season. The resulting vocal delivery proved remarkably accurate, successfully avoiding any minor hoarseness present in the original reference recording.

Exposing the personal avatar to non-tech-savvy friends yielded mixed reactions, described by observers as both fascinating and eerie.

Testing the interactive avatar revealed a more rigid operational framework. Because the model is deterministic—meaning it is strictly bound to the parameters of its training material—it repeatedly redirected off-topic questions back to the venture fraud story. Attempts to prompt the avatar about personal trivia, such as previous career history or residential neighborhoods in New York, were met with polite redirections to the designated article.

While friends noted that the synthesized voice and visual likeness in the interactive model felt slightly less precise than the personal script-reading avatar, the overall resemblance remained close enough to elicit a subtle sense of unease. Family members offered similarly enthusiastic feedback, with relatives attempting to trick the system by asking private questions only close family members would know. When the model repeatedly deflected these personal queries in favor of discussing venture capital fraud, family members joked about the surreal nature of interacting with a digital copy.

Experiences of this nature inevitably prompt broader questions regarding the future trajectory of journalism and media consumption. The prospect of daily news broadcasts hosted entirely by virtual avatars raises fundamental questions about public reception. While financial investors express immediate skepticism regarding audience acceptance, the ongoing proliferation of AI-generated content across social media and digital publishing platforms demonstrates a shifting landscape. The debate persists over whether digital twins will eventually augment or replace human journalists, or whether corporate executives might eventually prefer addressing an AI avatar over a human reporter.

At its core, the most rewarding aspect of journalism involves interpersonal connection, rigorous reporting, and deep exploration of complex topics. The foundational currency of journalism remains trust, an element that cannot easily be outsourced to automated systems.

Beyond the realm of media, the commercial appeal of personal digital cloning is evident. The ability to delegate routine professional obligations, such as catching up on correspondence following an extended vacation, suggests a future where a functional digital proxy remains perpetually accessible to address inquiries.

As corporate adoption of avatar technology continues to evolve, the personal experience of managing a digital twin leaves lingering, complex impressions. Moving past the initial novelty reveals a quiet realization when observing the silent avatar waiting for further interaction—a momentary expectation for an involuntary blink, a spontaneous smile, or some subtle confirmation of conscious awareness.

Because these deterministic digital twins operate within strict algorithmic boundaries, those spontaneous human behaviors will never materialize. However, the technology clearly demonstrates how easily individuals could develop deep psychological attachments to non-deterministic systems powered by open-ended conversational chatbots.

Reflecting on these developments, younger generations may take considerable time adjusting to the pervasive presence of digital twins, which continue to blur the line between science fiction and everyday reality. Ultimately, digital avatars often feel less jarring than physical humanoid robotics; should interactions become overly complex or bizarre, the simple solution remains the ability to log off and walk away.

Leave a Reply

Your email address will not be published. Required fields are marked *