
Convai vs Inworld vs DIY: How to Add Talking NPCs to Your Game in 2026
Convai vs Inworld vs DIY: How to Add Talking NPCs to Your Game in 2026
You want a Phaser or JavaScript NPC that talks back. Not a cutscene voice line and not a branching dialogue tree, but a character that hears the player, reasons about what was said, and answers out loud. In 2026 you have three real paths: a hosted character platform, voice and inference components, or a speech loop you build in the browser. The thesis of this article is simple: this is a “who owns the character” decision, and every pricing and portability question follows from it.
Scope note: this comparison is based on official documentation and pricing pages read on 2026-10-04, and no tool was run hands-on.
What Changed in 2026
The biggest shift is that Inworld retired its hosted Character Studio between March and September 2025 and repositioned as a voice and inference provider, so the old “pick a hosted character engine” market is now a choice between one hosted platform and two build-it paths.
Convai, by contrast, still runs the full hosted pipeline: speech recognition, LLM reasoning, text-to-speech, text-to-action, and lip sync, plus a per-character knowledge bank, long-term memory, Vision, and Actions, per Convai’s documentation.
The browser half of the story has not changed much. The Web Speech API still offers a speech-to-text half and a text-to-speech half, and the DIY loop of recognition, LLM call, and synthesis remains a viable weekend project for a Phaser game.
If you come from Unity, one caution: Unity’s 2026 AI tools are editor-side only, an in-editor assistant, an AI gateway, an MCP server, and an official plugin for Claude Code, Codex, and Grok, they require Unity 6 or later, they are not an NPC runtime, and Unity’s FAQ states that Unity Muse is a deprecated product offering, per Unity’s AI features page and Unity’s blog.
The 3-Way Comparison at a Glance
Convai sells a hosted character platform where Convai hosts the character and knowledge. Inworld sells voice and inference components where you build the character and own memory and actions. The DIY browser path is your own speech-to-LLM-to-speech loop where you own everything and pay only LLM token cost.
| Dimension | Convai | Inworld | DIY Browser |
|---|---|---|---|
| What you buy | Hosted character platform | Voice and inference components | Your own speech-to-LLM-to-speech loop |
| Who owns the character | Convai hosts the character and knowledge | Developer builds the character and owns memory and actions | Developer owns everything |
| Pricing model | Monthly tiers; interactions reset monthly | Credit/usage model; paid credits roll over up to 3 months | LLM token cost only |
| Browser / Phaser fit | WebGL: full for core NPC features | API/SDK-based; not a hosted character engine | Native Web Speech API path; SpeechRecognition limited |
| Beginner lift | Lowest for a ready NPC | Medium: more integration work | Highest: you build and debug the loop |
Convai: The Hosted Character Platform
Convai is the closest thing to a one-stop shop for a talking NPC: speech recognition, LLM reasoning, text-to-speech, text-to-action, and lip sync, plus a per-character knowledge bank, long-term memory, Vision for camera and scene perception, and Actions, all hosted for you.
Pricing is monthly, with yearly discounts in brackets, on Convai’s pricing page. Free is $0 with 100 interactions per month, 1 character session concurrency, and 60 minutes of Cloud Avatar Studio and Convai Sim. Indie Dev is $29/mo ($22/yr) with 3,000 interactions, 1 concurrency, 600 minutes, and 5 MB of knowledge. Professional is $99/mo ($69/yr) with 10,000 interactions, 3 concurrency, 2,400 minutes, and 20 MB. Scale is $499/mo ($299/yr) with 50,000 interactions, 15 concurrency, 10,200 minutes, and 100 MB. Business is $1,199/mo ($499/yr) with 125,000 interactions, 30 concurrency, 21,000 minutes, and 300 MB. Enterprise is custom.
An “interaction” is one user input plus one character response, and quotas reset monthly and do not roll over, per the pricing page. Knowledge Bank and long-term memory are not on the Free tier, and the monthly active end users allowed per tier are 1, 75, 250, 1,250, and 2,500 respectively.
For a browser game, the WebGL support matrix matters. Voice conversation, lip sync, actions, dynamic context, emotion, long-term memory, and Vision via canvas capture are all fully supported, while spatial audio, screen share, microphone device selection, and microphone test are not supported, per Convai’s WebGL platform guide. It requires an HTTPS origin, with localhost exempt, and a user gesture before audio starts. Engine support covers the Unity SDK, the Unreal Engine plugin, and web plugins, per the documentation. Enterprise adds on-prem or private deployment, SLAs, and data ownership, per Convai’s enterprise page.
Inworld: Voice and Inference Components
Inworld no longer runs a hosted character engine. It now provides first-party TTS and STT models, the LLMs you choose, and the inference underneath, while you build the character in your own engine or backend and own its memory, knowledge, and allowed actions.
The pivot is documented on Inworld’s blog: Character Studio was retired between March and September 2025, the old studio host now redirects to the new developer platform, and the features moved to the Inworld Agent Runtime, where the developer builds the character in their own engine or backend. The current positioning on Inworld’s homepage is “a research lab and inference provider for realtime AI at consumer scale.”
The product lineup is Realtime TTS, Realtime STT, a Realtime API that combines STT, LLM, and TTS in one WebSocket session and is OpenAI Realtime protocol compatible, Realtime Inference, a Realtime Router that routes to 220+ LLMs at zero markup, and Compute, per the homepage.
Pricing on Inworld’s pricing page runs On-Demand free with up to 70 minutes of TTS or 400 minutes of STT free, then Creator at $25/mo, Builder at $100/mo, Developer at $300/mo, Growth at $1,500/mo, and Enterprise custom. Per-1M-character Realtime TTS-2 rates fall from $25 on On-Demand to $12.50 on Growth and as low as $5 on Enterprise, TTS-2 Flash runs $15 down to $7, and STT runs $0.15/hr down to $0.10/hr. LLMs are billed at provider cost with 0% markup, and credits are dollar-denominated and roll over up to 3 months on a paid plan.
Engine support includes the Unreal AI Runtime SDK generally available, the Unity AI Runtime SDK in early access, a Node.js SDK, a visual graph editor, and pre-built templates including Character, Metahuman, and Lipsync, per Inworld’s SDK announcement. For a Phaser game, the Node.js SDK and the Realtime API over WebSocket are the natural fit, but you wire recognition, reasoning, and speech into your game loop yourself.
DIY Browser: Roll Your Own Talking NPC
The DIY path uses the Web Speech API: SpeechRecognition turns player speech into text, you call your own LLM with the character’s system prompt and history, and SpeechSynthesis speaks the reply. You pay only the LLM token cost and own all state.
The two halves have very different maturity. SpeechSynthesis is Baseline and widely available since September 2018, per MDN. SpeechRecognition has limited availability: it is not Baseline, and on some browsers such as Chrome the audio is sent to a server-based recognition engine, so it does not work offline, per MDN.
A minimal loop looks like this:
const recognition = new (window.SpeechRecognition || window.webkitSpeechRecognition)();
recognition.lang = "en-US";
recognition.onresult = async (event) => {
const transcript = event.results[0][0].transcript;
const reply = await getNpcReply(transcript); // your LLM fetch with system prompt + history
speechSynthesis.speak(new SpeechSynthesisUtterance(reply));
};
recognition.start();
The tradeoffs are honest ones. There is no per-interaction vendor fee and no hosted character backend, and the character’s memory, knowledge, and allowed actions live in your code. In return, you build and debug the loop yourself, you handle the HTTPS plus user-gesture requirement before audio starts, and you accept that SpeechRecognition is not supported in Firefox for Android, per MDN.
Which Should You Pick?
If you want a ready NPC fast and accept that the character lives on someone else’s platform, pick Convai. If you want professional voice and inference components while owning the character, pick Inworld. If you want full ownership and minimal cost, build the DIY loop.
A few rules of thumb. If you are shipping a small Phaser game this month and want voice, lip sync, and memory without writing a pipeline, Convai’s Free or Indie Dev tier is the lowest lift, and its WebGL support covers the core NPC features. If you already have a backend and want first-party TTS and STT with the LLMs you choose, Inworld’s Realtime API and Router fit, and the character belongs to your game. If you are prototyping, learning, or want zero per-interaction cost, the DIY loop is the cheapest and most educational path.
Remember the ownership question as you plan to grow. Convai’s quotas reset monthly and do not roll over, Inworld’s paid credits roll over up to 3 months, and the DIY loop has no quota at all, only your LLM bill. Pick the model that matches where you want your character to live in two years.
FAQ
These are the questions beginners ask most when adding a talking NPC to a Phaser or JavaScript browser game. Each answer assumes you care about who owns the character, what it costs, and how much of the pipeline you want to build yourself.
Can I use Convai with a Phaser browser game?
Yes, through Convai’s web plugins, with WebGL support that is full for voice conversation, lip sync, actions, dynamic context, emotion, long-term memory, and Vision via canvas capture, per Convai’s WebGL guide. You need an HTTPS origin, with localhost exempt, and a user gesture before audio starts.
Does Inworld still offer a hosted character engine?
No. Inworld retired its Character Studio between March and September 2025 and now provides voice and inference components, with the character built in your own engine or backend, per Inworld’s blog. The old studio host redirects to the new developer platform.
Is the Web Speech API enough for a shipping game?
For a small or experimental game, yes, with caveats. SpeechSynthesis is Baseline and widely available, but SpeechRecognition is not Baseline, does not work offline on some browsers such as Chrome, and is not supported in Firefox for Android, per MDN. You also handle the LLM call, the system prompt, and the history yourself.
In 2026, adding a talking NPC to a Phaser or JavaScript game comes down to one question: who owns the character? Convai hosts the character and gives you the fastest path to a ready NPC. Inworld gives you professional voice and inference components while you own the memory and actions. The DIY browser loop costs only LLM tokens and gives you total ownership, plus total responsibility. Choose the ownership model you can live with, and the rest is integration work.