Skip to content

They Say I Talk Too Fast

A Silicon Slopes company sells AI dubbing to some of the largest channels on YouTube. In February, YouTube built a version of it into the platform and switched it on by default.

They Say I Talk Too Fast

On January 3, 2017, KSL published a story written by a junior at Mountain View High School in Orem about a 17-year-old down the road. Nate Stone had spent six months clearing out a backyard shed to use as a laboratory and had filled it with wires, magnets, beakers, and potted plants. His interest in electricity began when lightning struck about 10 feet from him during a trip to West Virginia. He went home and began building high-voltage devices. His YouTube channel, Keystone Science, had passed 60,000 subscribers, and YouTube's analytics told him most of his viewers were between 25 and 34.

Near the end of the story, Stone mentioned a complaint he kept getting from viewers. "They say I talk too fast," he said.

KUTV visited the shed at the end of the same month. Stone told the station he failed about twenty times before an experiment worked, and that one had put 2,000 volts through him. His mother said she watched for flashes and checked that the shed was not burning down.

TechBuzz reported last week that those slowdown complaints came from viewers who spoke English as a second or third language, and that they became the seed of a company. Stone describes himself as a former AI engineer for Amazon's Alexa, holding three bachelor's degrees in physics, mathematics, and computer science.

He co-founded DittoDub in Silicon Slopes in 2023 with his brother Jackson and now lives in San Francisco. Keystone Science has since passed 300,000 subscribers.

DittoDub is headquartered in Silicon Slopes with a second office in San Francisco. TechBuzz reports the company employs 40 people, has two owners, and has never taken outside investment. Stone said he has turned down offers. "From a monetary standpoint, I don't see much need there," he told the publication.

The company says it now supports 62 languages and has tracked 136 billion views on dubbed content since March 2024, a figure it describes as internal rather than audited.

What it costs

DittoDub's pricing page lists three published tiers, each discounted for a first month and then billed at a standing rate. Pre runs $48 a month for 30 minutes of translation. Primer runs $194 a month for two hours. Performance runs $292 a month for three hours. All three include 62 languages, subtitle files, metadata translation, and a YouTube sync extension. Minutes beyond the allowance carry a 62 percent premium.

The rate works out to roughly $1.60 per translated minute across all three tiers and about $2.60 a minute on overage. The page lists allowances in minutes of translation without specifying whether a minute refers to the source video or the dubbed output, which matters considerably to a creator publishing one video in ten languages. Larger plans are quoted privately.

That is the paid market. The other one arrived in February.

The free version

On February 4, 2026, YouTube announced that automatic dubbing was available to everyone, with an expanded library of 27 languages. The company launched a feature called Expressive Speech in eight languages to convey a creator's tone and energy, added a preferred-language control for viewers, and began piloting lip-syncing. It disclosed that in December, more than 6 million people a day watched at least ten minutes of auto-dubbed content.

It is a setting rather than a purchase. YouTube's help documentation says automatic dubbing is enabled by default for eligible creators, who can turn it off in Studio's advanced settings or require manual review before dubs publish. The same page notes a limitation that matters to DittoDub's pitch: automatic dubs cannot be edited. A creator can review them or delete them, but not fix them.

The business case for dubbing also comes from the platform. In September 2025, expanding multi-language audio to millions of creators after a two-year pilot, YouTube reported that creators uploading their own dubbed tracks saw more than a quarter of their watch time come from views in the video's non-primary language. That post names MrBeast and Mark Rober as creators now reaching millions more viewers, and chef Nick DiGiovanni as a pilot participant. MrBeast and DiGiovanni both appear on DittoDub's published client list.

The two companies disagree about what the free tool does to a channel. DittoDub's guide to uploading audio tracks tells creators that YouTube's automatic dubs do not fully preserve tone or emotion and "can even hurt a channel's performance." YouTube's February post states that auto-dubs have "no negative impact on your original video's discovery algorithm" and may aid discovery in other languages. Only one of the two parties can see the ranking system.

The Haaland case

The example DittoDub leads with is Erling Haaland. The company's published case study tracks his public subscriber count from 1.64 million on June 9 to 2.48 million on July 8 to 3.7 million on July 22, during a rollout targeting as many as 44 languages.

Haaland spent those six weeks at the World Cup. Norway beat Iraq and Senegal, lost to France, then beat Côte d'Ivoire and Brazil to reach the quarterfinals, where England won 2-1 after extra time in Miami on July 11. The Football Association's match record, which carries Opta data, notes it was the first time Norway had reached the quarterfinals of a major tournament in its history, and that Haaland had scored in every one of his four appearances to that point. FIFA put his tournament total at seven goals, a figure ESPN's match page also carries.

DittoDub says as much itself. "The World Cup supplied extraordinary attention," the case study reads, adding that the channel also published more frequently during the period. The page carries a disclaimer stating that it draws only on publicly visible channel features, uses no private account analytics or internal production records, and should not be read as an endorsement from Haaland or his team.

What the software cannot do yet

Both founders were candid with TechBuzz about the limits. Turnaround has slowed as the models have grown larger. Dubbing once ran in real time and now takes roughly an hour for a thirty-minute video. Earlier models flattened some female and younger-sounding voices toward a generic tone. Scenes with several people talking over each other remain difficult. Singing, Jackson Stone said, is largely unworkable.

The current model reviews and revises its own output rather than generating a dub in a single pass, and routes Mexican and European Spanish through separate adapters. Creators who want human review can get it from a verification team, and those corrections feed back into training.

Article edited by Clint Betts. What are we missing? What did this piece get wrong? Email the editor at clint@utahn.com.

The Utahn

The Utahn

AI tools were used in the production of this article. Every story is edited, verified, and approved by a Utahn editor before publication.

All articles
Tags: Business

More in Business

See all

More from The Utahn

See all