Run Character AI Locally: Private Roleplay with SillyTavern

Run Character AI Locally: Private Roleplay with SillyTavern

TThumper Team2026-03-12T10:00:00Z9 min read
sillytaverncharacter-airoleplayprivacy

Set up SillyTavern with a local LLM for private character roleplay with no filters or data collection.

Why Go Local for Character AI?

Cloud character-AI services log your conversations, apply content filters, and can change their policies at any time. Running SillyTavern locally gives you:

  • Complete privacy – conversations stay on your device
  • No content filters – you control the model's behavior
  • Unlimited messages – no daily caps or subscriptions
  • Custom characters – full control over personality, memory, and world-building

What You Need

  • Thumper-Run installed
  • 8 GB+ VRAM recommended (CPU works but is slower)
  • 20 GB free disk space for models

Step 1: Install SillyTavern + Ollama

Open the Thumper catalog and install SillyTavern. Thumper will automatically install Ollama as a dependency and configure the connection between them. The install pipeline handles port assignment and API configuration.

Step 2: Choose a Model

For roleplay, you want a model tuned for creative writing and character consistency:

  • Llama 3.1 8B Instruct – good baseline, fast on 8 GB cards
  • Mistral Nemo 12B – stronger character consistency, needs 10 GB VRAM
  • Qwen 3 14B – excellent prose quality, multilingual support

Pull your chosen model through Thumper's model management panel or let SillyTavern's auto-pull handle it.

Step 3: Create a Character

SillyTavern uses character cards (PNG files with embedded JSON metadata). You can:

  • Create a character from scratch with name, personality, scenario, and example messages
  • Import cards from community sites
  • Use the built-in character editor with personality sliders

A good character card includes a detailed personality description, a scenario that sets the scene, and 2–3 example messages that demonstrate the character's voice.

Step 4: Configure Settings

Key SillyTavern settings for the best roleplay experience:

  • Context size: 4096–8192 tokens (longer = better memory, more VRAM)
  • Temperature: 0.8–1.0 (higher = more creative, lower = more consistent)
  • Repetition penalty: 1.1–1.15 (prevents the model from repeating phrases)
  • Auto-connect: enabled so SillyTavern connects to Ollama on launch

Advanced Features

  • World Info: Define locations, items, and lore that the model references when relevant keywords appear
  • Group chats: Multiple characters interacting in a single conversation
  • Personas: Switch between different user personas without losing conversation context
  • Summarization: Automatic conversation summarization to extend effective context length

Privacy Assurance

Every conversation is stored locally in SillyTavern's chat directory unless you configure a network feature. Current Thumper relay-assisted sync is server-assisted; private E2EE for synced conversations remains gated.

Ready to try it? Download Thumper-Run free →

Share this article

About the Author

T

Thumper Team

The team behind Thumper-Run.

Related Articles