TotalApp lets you pick the engine behind every AI Writer, Assistant, and Semantic Search feature — from fully local Ollama models to in-browser inference to any hosted cloud provider — all from one Settings screen.
Every AI Writer, AI Assistant, and Semantic Search screen in TotalApp routes through the same Agentic configuration. Pick a Writer Engine for text generation and, independently, an Embedding Engine for knowledge ranking — both from Settings → Agentic.
Run text generation and embeddings entirely on your own machine with Ollama — no traffic ever reaches TotalApp's servers. Download models like mistral, llama3.3, deepseek-r1, or qwen2.5, grouped by whether they're built for everyday chat, heavy reasoning, or lightweight low-power tasks.
In-Browser LLM (WebGPU) runs models like SmolLM2, TinyLlama, and Qwen2.5-Coder-1.5B directly inside the browser tab's own memory and GPU — nothing to install outside the browser, and the model downloads once and stays cached.
Bring your own API key for Anthropic, OpenAI, Google AI Studio, DeepSeek, Cohere, Groq, or OpenRouter. Providers are grouped by architecture — Direct, Proxy/Router & Serverless, or Aggregator — so you always know exactly where your requests go.
Local for privacy, in-browser for zero setup, or cloud for the strongest models — TotalApp's Agentic settings let you choose, and change your mind anytime.