One model talks, another moves
A character in the browser where the voice and the body are different models, and the film says which voice was scripted.
- Category
- AI
- Who it is for
- Someone who wants a character they can talk to in the browser, and will bring keys.
- Possible model
- Open source
- Difficulty
- Hard
- Inspired by
- open-annie · open-annie
- Original launch
- Watch the launch · Oct 7, 2026
Difficulty is an editorial estimate of technical and operational complexity, not a score.
Startup idea analysis by Startup Ideas. Last reviewed Oct 7, 2026. How we read a launch.
The startup idea
Rehan Sheikh, whose bio says he is an engineer at fal, posted open-annie. The repo is github.com/rehan-remade/open-annie. The readme says one model talks and another moves. GPT-Live-1 speaks over WebRTC. Jev reads the transcript and picks the face, the gesture, and the body. The post says that pick is about 120 ms. The mouth is lip-synced from the audio. The picture runs in the browser. The server is a broker on fal. The readme says you can look around with no keys. Talking needs an OpenAI key with GPT-Live access and a TypeSafe key. It also says the film is one recorded session: Annie's lines are the model's, and the other voice is six scripted lines from a synthetic microphone. The license file is Apache 2.0. This page did not talk to her.
The opportunity is the split, and the label on the film. A character that talks and moves is easy to fake if both halves are pre-recorded. Here the readme says which half was scripted. A smaller product is one character, a look-around that needs no key, a live mode that asks for keys, and a line on the film when a voice is not a person. If you hide the script, you do not have this product.
Do not say this page measured 120 ms. The post says about 120 ms. The readme's timing is for that recorded session. Do not invent a price. The keys are the customer's. Do not treat Annie's lines as a support agent for a company.
The problem
A talking head is one model doing two jobs, and the demo hides which half was recorded.
Who would pay for this?
- Primary
- Someone who wants a character they can talk to in the browser, and will bring keys.
- Secondary
- A developer who wants the split, voice in one place and motion in another, and can read the repo.
- Early adopter
- One avatar, the no-key preview, then one live sentence.
How could it make money?
Possible models. Nothing here is a claim about what open-annie charges.
- No price on the post. Apache 2.0 is the license.
- Live talk spends the customer's OpenAI and TypeSafe keys. Say that before the microphone.
- A hosted version with your keys would be a different product.
What would the MVP look like?
- One character in the browser.
- A preview that needs no key.
- Voice from one model, motion from another.
- A line that the other voice in the film is scripted.
- The keys named before the live mode.
How to validate the idea
- Ask five people if they can tell the scripted voice from the live one after you label it.
- If they only want the preview, the keys were the wall.
- Do not add a second character until one live sentence works without you in the room.
How to differentiate
- Two models, two jobs.
- The scripted half is written down.
- Look around before you pay a key.
- Do not become a companion app. The split is the product.
Why this could be interesting now
The post puts a talking character in the browser and the readme says which voice was performed. The opening is that label, and a preview that does not need a key.
Original product launch
This launch is inspiration and evidence that someone shipped, not a partnership. The idea above is our reading of the broader opportunity. open-annie did not write it.
- Product
- open-annie
- Source
- open-annie
- Film
- open-annie, one model talks and another moves
- Posted
- Oct 7, 2026 · 1:37
Product that inspired this idea
Rehan, who works at fal, open-sourced a character in the browser. GPT-Live talks. Jev picks the face and the gesture. The post says that pick is about 120 ms. The readme says the other voice in the film is scripted, talking needs your own keys, and the license is Apache 2.0.
Posted by @rehan_shei. Watch the original launch. No separate product site is listed unless the post itself is the source.
- voice
- character
Related startup ideas
- A coworker that stays on after the chat
An agent with its own computer that keeps a project page, and can be reached from the chat app a team already uses.
- A self-hosted assistant for personal apps
An assistant a person can host, pointed at the apps they already use, including computer use.
- A worker that lives in a folder
Give each worker a folder on the machine, a department, and a job, and say on the first screen whether it asks before it acts.
- Ask the computer, then approve the change
Put the agent on the desktop the person already uses, and stop it before it changes a file.
- Memory that survives the next coding session
Sell a memory layer that learns how one developer works, then carries that into the next agent run.
- Robot control from video, as a research product
A path from video models to robot control, sold first as weights and a narrow demo, not a factory.
Looking for more ideas?
The full list is startup ideas for 2026, each one tied to a launch you can watch.
Browse startup ideas for 2026