Are you the author? Sign in to claim
๐ญ AI Avatar / digital human platform โ upload a photo, clone a voice, talk to any face in real time with lip-sync video
Upload a photo ยท Clone a voice ยท Talk to any face in real time
Quick Start ยท Features ยท Architecture ยท GPU / AWS Deploy ยท API ยท Roadmap
The most complete open-source AI avatar / digital human system. Real-time talking-head lip-sync ยท Zero-shot voice cloning ยท Multi-LLM ยท Runs 100% locally or on AWS.
AvatarAI is an open-source, production-ready platform for building photorealistic AI avatar conversations. Upload any face photo, clone a voice from a 5-second audio clip, and have a real-time conversation โ with lip-sync video generated on every single response.
[mic] โ Whisper STT โ Claude / GPT / Ollama (streaming) โ Chatterbox TTS โ MuseTalk lip-sync โ [video]
< 2โ4 s to first video chunk on AWS GPU >
What makes AvatarAI different:
g5.xlarge for true real-time (~30 FPS)| AvatarAI | Duix-Avatar | Linly-Talker | AIAvatarKit | |
|---|---|---|---|---|
| Real-time conversation | โ WebSocket streaming | โ offline video gen | โ (Gradio / WebRTC spin-off) | โ |
| Lip-sync video | โ MuseTalk V1.5 | โ proprietary models | โ multiple engines | โ (drives external avatars) |
| Voice cloning | โ 10 s, 23 languages | โ | โ | โ |
| Barge-in / interruption | โ | โ | โ (stream variant) | โ |
| Local / free LLM | โ Ollama, vLLM | โ | โ | โ |
| Web app with auth & history | โ Next.js + JWT + Postgres | โ Windows client | โ Gradio demo UI | โ library |
| Rate limiting, CI, tests, IaC | โ | โ | โ | โ |
| License | MIT | custom | MIT | Apache-2.0 |
Toolkits like Linly-Talker are great research playgrounds; Duix ships a Windows product. AvatarAI is the one you can deploy as a real multi-user web service.
| Category | Details |
|---|---|
| ๐ค LLM Backends | Claude (prompt-cached) ยท GPT-4o ยท Ollama / vLLM / LM Studio (local, free) |
| ๐ค Voice Cloning | Record 10โ60 s โ Chatterbox Multilingual zero-shot cloning |
| ๐ฃ๏ธ Speech-to-Text | Whisper (faster-whisper, CUDA), decodes browser WebM natively |
| ๐ฌ Lip-Sync Video | MuseTalk V1.5 persistent worker (30 FPS on GPU) ยท FFmpeg fallback (CPU) |
| โก Streaming Pipeline | Live LLM tokens + per-sentence video chunks over WebSocket |
| โ Barge-In | Speak or hit stop mid-reply โ in-flight turn cancels in ms |
| ๐ TTS Fallback Chain | chatterbox โ edge-tts (free neural voices) โ gTTS โ never silent |
| ๐ Emotion Detection | Live emotion badges per message |
| ๐ 23 Languages | Whisper multilingual STT + Chatterbox multilingual TTS |
| ๐ Local-First Storage | USE_LOCAL_STORAGE=true โ no AWS needed for dev |
| ๐ Auth & Sessions | JWT authentication, conversation history, persistent sessions |
| ๐ Observability | Prometheus ยท Celery Flower ยท Sentry ยท structured logging |
| ๐งช Tested | Full pytest suite โ users, avatars, sessions, health checks |
| ๐ AWS GPU Deploy | One-command g5.xlarge deploy with CUDA 11.8 + float16 |
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Browser / Client โ
โ โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โAvatar Studioโ โ Voice Studio โ โ Chat Interface โ โ
โ โ (upload) โ โ (cloning) โ โ Idle anim + chunks โ โ
โ โโโโโโโโฌโโโโโโโ โโโโโโโโฌโโโโโโโโ โโโโโโโโโโโโฌโโโโโโโโโโโโ โ
โโโโโโโโโโโผโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโ
โ REST โ REST โ WebSocket
โผ โผ โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ FastAPI Backend โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ WebSocket Manager โ โ
โ โ split sentences โ TTS โ MuseTalk โ stream chunks โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โโโโโโโโโโโโ โโโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโโโโโโ โ
โ โ Whisper โ โClaude/GPT โ โ XTTS v2 โ โ MuseTalk โ โ
โ โ STT โ โ / Llama โ โ TTS โ โ (GPU/CPU) โ โ
โ โโโโโโโโโโโโ โโโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโโโโโโ โ
โ โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโโโโโโ โ
โ โPostgreSQLโ โ Redis โ โ Celery โ โ Local FS / S3 โ โ
โ โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโโโโโโ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
[User types / speaks]
โ
โผ
Whisper STT โโโโโโโโโโโโโโโโโโบ transcript
โ
โผ
Claude / GPT / Llama โโโโโโโโโบ full response text
โ
โผ
Split into sentences โโโโโโโโโบ ["Hello!", "How are you?", ...]
โ
โโโ sentence 1 โ XTTS โ MuseTalk โ video_chunk WS โ browser plays
โโโ sentence 2 โ XTTS โ MuseTalk โ video_chunk WS โ queued
โโโ sentence N โ XTTS โ MuseTalk โ video_chunk WS โ queued
ai-avatar-system/
โโโ backend/ # FastAPI application
โ โโโ app/
โ โ โโโ api/v1/ # REST endpoints (users, avatars, sessions, messages)
โ โ โโโ services/ # Core services (LLM, TTS, STT, animator, storage)
โ โ โโโ models/ # SQLAlchemy DB models
โ โ โโโ websocket.py # Real-time WebSocket handler + sentence streaming
โ โโโ alembic/ # Database migrations
โ โโโ models/MuseTalk/ # MuseTalk V1.5 (lip-sync engine)
โ โ โโโ scripts/
โ โ โโโ musetalk_worker.py # Persistent worker (models loaded once)
โ โโโ tests/ # pytest suite
โ โโโ Dockerfile # CUDA 11.8 base image
โ โโโ requirements.txt
โโโ frontend/ # Next.js 14 application
โ โโโ app/ # App Router pages
โ โโโ components/ # React components (ChatInterface, IdleAvatar, etc.)
โ โโโ lib/api.ts # Axios API client
โ โโโ store/ # Zustand global state
โโโ nginx/
โ โโโ nginx.conf # Reverse proxy (HTTP โ backend/frontend, WebSocket)
โโโ infrastructure/
โ โโโ main.tf # AWS Terraform (ECS, RDS, ElastiCache, S3, CloudFront)
โ โโโ variables.tf
โโโ scripts/
โ โโโ setup_musetalk.sh # Download MuseTalk models (~9 GB)
โ โโโ deploy-aws.sh # One-command EC2 GPU deployment
โโโ docker-compose.yml # Development (CPU) โ all services
โโโ docker-compose.prod.yml # Production overrides (GPU, no bind mounts, logging)
โโโ deploy.sh # ECR push + Terraform deploy (ECS path)
โโโ .env.example # Development env template
โโโ .env.prod.example # Production env template
git clone https://github.com/PunithVT/ai-avatar-system.git
cd ai-avatar-system
cp .env.example .env # add your ANTHROPIC_API_KEY (or OPENAI_API_KEY)
docker compose up -d
| Service | URL |
|---|---|
| ๐ฅ๏ธ Frontend | http://localhost:3000 |
| โ๏ธ Backend API | http://localhost:8000 |
| ๐ Swagger Docs | http://localhost:8000/docs |
| ๐ธ Celery Flower | http://localhost:5555 |
No AWS required. Set
USE_LOCAL_STORAGE=true(default) โ uploads saved tobackend/uploads/.
Want something to talk to immediately? Seed three ready-made demo avatars (AI-generated faces + personalities):
backend/venv/bin/python scripts/seed_demo.py # or any python with `requests`
backend/venv/bin/python scripts/seed_demo.py --with-voices # + cloned demo voices
Prebuilt images are also published on every release โ ghcr.io/punithvt/ai-avatar-system-backend and โฆ-frontend.
# Backend
cd backend
python -m venv venv && source venv/bin/activate
pip install -r requirements.txt
cp ../.env.example ../.env
alembic upgrade head
uvicorn main:app --reload --port 8000
# Frontend (new terminal)
cd frontend
npm install
npm run dev
# Download models (~9 GB, one-time)
bash scripts/setup_musetalk.sh
# Set in .env
AVATAR_ENGINE=musetalk
# Restart
docker compose restart backend
MuseTalk achieves 30 FPS at 256ร256 on a V100-class GPU (source: MuseTalk paper). On CPU it is 30โ50ร slower. Deploying on AWS gets you genuine real-time performance.
| Instance | GPU | VRAM | Spot $/hr | MuseTalk FPS |
|---|---|---|---|---|
g4dn.xlarge | T4 | 16 GB | ~$0.16 | ~15โ20 FPS |
g5.xlarge | A10G | 24 GB | ~$0.30 | ~30 FPS โ |
g6.xlarge | L4 | 24 GB | ~$0.24 | ~30 FPS โ |
Recommended: g5.xlarge Spot (~$72/mo at 8 hrs/day).
# 1. Launch g5.xlarge with Ubuntu 22.04 LTS, SSH in, then:
bash <(curl -fsSL https://raw.githubusercontent.com/PunithVT/ai-avatar-system/main/scripts/deploy-aws.sh)
# 2. Fill in API keys:
nano /opt/ai-avatar-system/.env.prod
# 3. Redeploy with your keys:
bash /opt/ai-avatar-system/scripts/deploy-aws.sh --update
The script automatically:
cp .env.prod.example .env.prod # fill in your values
docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d
What docker-compose.prod.yml adds over development:
nvidia driver, count=1) for backend + celery-workerfloat16 inference enabled automatically on CUDA โ ~2ร speedupmusetalk_models volume (survive container restarts)# Check GPU is visible in container
docker exec avatar-backend python -c "
import torch
print('CUDA:', torch.cuda.is_available())
print('GPU:', torch.cuda.get_device_name(0))
print('VRAM:', round(torch.cuda.get_device_properties(0).total_memory/1024**3,1), 'GB')
"
# Expected on g5.xlarge:
# CUDA: True
# GPU: NVIDIA A10G
# VRAM: 24.0 GB
# Live GPU utilisation
docker exec avatar-backend nvidia-smi
For a fully managed ECS deployment with RDS + ElastiCache + CloudFront:
cd infrastructure
terraform init
terraform apply -var="environment=production"
bash deploy.sh production
Powered by Chatterbox Multilingual (Resemble AI) โ zero-shot voice cloning from a 10-second sample, in 23 languages.
Every TTS response then uses your cloned voice.
# REST API
curl -X POST http://localhost:8000/api/v1/voices/clone \
-F "audio=@my_voice.wav" -F "name=My Voice" -F "language=en"
POST /api/v1/users/register { "email": "...", "username": "...", "password": "..." }
POST /api/v1/users/login form: username=... password=... โ { "access_token": "..." }
# All protected routes:
Authorization: Bearer <access_token>
POST /api/v1/avatars/upload Upload photo (multipart: file + name)
GET /api/v1/avatars/ List avatars
DELETE /api/v1/avatars/{id} Delete avatar
PUT /api/v1/avatars/{id}/voice Assign voice to avatar
POST /api/v1/sessions/create { "avatar_id": "..." }
POST /api/v1/sessions/{id}/end
GET /api/v1/messages/session/{id}
WS /ws/session/{session_id}
Client โ Server:
{ "type": "text", "text": "Hello!" }
{ "type": "audio", "audio": "<base64-webm>" }
{ "type": "stop" } // barge-in: cancel the in-flight reply
{ "type": "set_voice", "voice_id": "<uuid>" } // attach a cloned voice (owner-checked)
{ "type": "set_language", "language": "es" }
{ "type": "ping" }
Server โ Client:
{ "type": "token", "token": "Hel" } // live LLM stream
{ "type": "transcription", "text": "Hello!" }
{ "type": "message", "content": "Hi!", "role": "assistant" }
{ "type": "video_chunk_start","total_chunks": -1 } // -1 = streaming, total unknown
{ "type": "video_chunk", "chunk_index": 0, "video_url": "...", "text": "Hi!" }
{ "type": "video_chunk_end", "sent_chunks": 3 }
{ "type": "status", "message": "Animatingโฆ", "stage": "animation" }
{ "type": "tts_fallback", "engine": "edge-tts", "voice_cloned": false, "message": "โฆ" }
{ "type": "interrupted", "message": "Previous response interrupted" }
{ "type": "error", "message": "Something went wrong" }
Key .env variables:
# LLM
LLM_PROVIDER=anthropic # anthropic | openai | ollama (local & free)
LLM_MODEL=claude-sonnet-4-6 # or gpt-4o ยท llama3.1 ยท qwen2.5 โฆ
ANTHROPIC_API_KEY=sk-ant-...
OPENAI_BASE_URL= # e.g. http://localhost:11434/v1 for Ollama / vLLM / LM Studio
# Avatar engine
AVATAR_ENGINE=musetalk # musetalk (GPU recommended) | simple (CPU fallback)
MUSETALK_PATH=models/MuseTalk
# TTS โ automatic fallback chain: chatterbox โ edge-tts โ gtts
TTS_PROVIDER=chatterbox
# STT
WHISPER_MODEL=large-v3-turbo # tiny | base | small | medium | large-v3 | large-v3-turbo
# Storage
USE_LOCAL_STORAGE=true # false โ AWS S3 (+ presigned URLs / CloudFront)
S3_BUCKET_NAME=...
# Auth (โฅ32 chars enforced at boot)
SECRET_KEY=$(python -c "import secrets; print(secrets.token_hex(32))")
JWT_SECRET_KEY=$(python -c "import secrets; print(secrets.token_hex(32))")
JWT_EXPIRATION_HOURS=24
| Library | Purpose |
|---|---|
| Next.js 14 + React 18 | App framework |
| TypeScript 5 | Type safety |
| Tailwind CSS | Styling |
| Zustand | Global state |
| Library | Purpose |
|---|---|
| FastAPI | Async REST API + WebSocket |
| SQLAlchemy 2 (async) | ORM with asyncpg |
| PostgreSQL 15 | Primary database |
| Alembic | Migrations |
| Redis 7 | Cache + Celery broker |
| Celery | Background tasks |
| Model | Purpose |
|---|---|
| Claude / GPT-4o / Ollama (local) | LLM conversation |
Whisper (faster-whisper) | Speech-to-text |
| Chatterbox Multilingual (Resemble AI) | TTS + zero-shot voice cloning, 23 languages |
| Edge TTS โ gTTS | Free no-GPU fallback voices |
| MuseTalk V1.5 | Photorealistic lip-sync (30 FPS on GPU) |
cd backend
pytest -v # all tests
pytest tests/test_health.py # single module
pytest --cov=app --cov-report=html # HTML coverage
No โ everything runs on CPU. MuseTalk takes 30โ90 s/sentence on CPU (the simple engine is instant). For real-time lip-sync, use an AWS g5.xlarge (~$0.30/hr spot) or any 16 GB+ NVIDIA card.
Yes โ set LLM_PROVIDER=ollama, run Ollama (ollama run llama3.1), and you have a fully local, free conversation stack: Whisper STT, local LLM, Chatterbox TTS, MuseTalk video.
Run python scripts/seed_demo.py โ it creates three demo avatars (AI-generated faces, distinct personalities) and optionally cloned demo voices with --with-voices.
Run bash scripts/setup_musetalk.sh โ downloads ~9 GB of models automatically.
The MuseTalk persistent worker loads all models into GPU VRAM on the first request (~60 s on GPU, ~5 min on CPU). Subsequent requests reuse the loaded models.
The pipeline degrades gracefully: chatterbox โ edge-tts (free Microsoft neural voices) โ gTTS. The UI shows a one-time notice when a cloned voice couldn't be applied.
A clear, well-lit frontal face photo (JPEG/PNG/WebP). Avoid sunglasses or heavy occlusion.
Contributions welcome! Read CONTRIBUTING.md before opening a PR.
git clone https://github.com/PunithVT/ai-avatar-system.git
git checkout -b feat/my-feature
# make changes + tests
git commit -m "feat(backend): add my feature"
git push origin feat/my-feature
MIT ยฉ 2026 โ see LICENSE for details.
โ ๏ธ Experimentelle Skill-Sammlung fรผr deutsches Recht (Arbeits-, Gesellschafts-, Insolvenz-, Datenschutz-, Prozessrecht u
Manage multiple Claude Code agents from TUI or Web with tmux and git worktrees
Project management using GitHub Issues + Git worktrees for parallel agent execution
Core skills library for Claude Code with 20+ battle-tested skills including TDD, debugging, and brainstorming