Agentic Solutions

Choose Where Your AI
Actually Runs.

TotalApp lets you pick the engine behind every AI Writer, Assistant, and Semantic Search feature — from fully local Ollama models to in-browser inference to any hosted cloud provider — all from one Settings screen.

Try Agentic Solutions See how it works
4
Writer engines supported
7+
Hosted cloud providers
0
Install for in-browser models
100%
Private & local with Ollama
How It Works

One setting, every AI feature

Every AI Writer, AI Assistant, and Semantic Search screen in TotalApp routes through the same Agentic configuration. Pick a Writer Engine for text generation and, independently, an Embedding Engine for knowledge ranking — both from Settings → Agentic.

  • Four Writer Engines: Ollama, In-Browser (WebGPU), Cloud LLM, Local CLI
  • Writer Engine and Embedding Engine are configured independently
  • Switch engines anytime — no migration, no lock-in to one provider
  • Mix a local embedding model with a cloud writer engine, or any combination
Configure Agentic Settings
totalapp.app/settings/agentic
Writer Engine
Ollama
Runs entirely on your machine
In-Browser LLM (WebGPU)
Runs in the browser tab, zero install
Cloud LLM (Hosted API)
Anthropic, OpenAI, Google & more
Local Models

Full privacy, full control

Run text generation and embeddings entirely on your own machine with Ollama — no traffic ever reaches TotalApp's servers. Download models like mistral, llama3.3, deepseek-r1, or qwen2.5, grouped by whether they're built for everyday chat, heavy reasoning, or lightweight low-power tasks.

  • Zero data leaves your device — the browser talks directly to your local Ollama instance
  • General Purpose, Reasoning & Heavy Tasks, and Lightweight & Micro model groups
  • Also powers local embeddings for Semantic Search ranking
  • No subscription or per-token cost — bring your own hardware
Read the Ollama Docs
127.0.0.1:11434
Local Ollama Models
Status: running on your machine — no cloud calls
mistral
general purpose
Installed
deepseek-r1
reasoning & code
Installed
llama3.2:1b
lightweight & micro
Available
In-Browser Models

Zero install, instant trial

In-Browser LLM (WebGPU) runs models like SmolLM2, TinyLlama, and Qwen2.5-Coder-1.5B directly inside the browser tab's own memory and GPU — nothing to install outside the browser, and the model downloads once and stays cached.

  • Powered by WebGPU, with automatic WASM fallback
  • Micro, Balanced, and High Performance model tiers to match your device
  • Best way to try AI Writer tools with no setup at all
  • Works fully offline once the model is cached by the browser
Read the In-Browser Docs
totalapp.app/settings/agentic
In-Browser LLM (WebGPU)
SmolLM2-135M
Micro — old or GPU-less machines
Llama-3.2-1B-Instruct
Balanced — everyday in-browser chat
Qwen2.5-Coder-1.5B
High performance — best code quality
Hosted Models

The strongest models available

Bring your own API key for Anthropic, OpenAI, Google AI Studio, DeepSeek, Cohere, Groq, or OpenRouter. Providers are grouped by architecture — Direct, Proxy/Router & Serverless, or Aggregator — so you always know exactly where your requests go.

  • Direct: Anthropic, OpenAI, Google AI Studio, DeepSeek
  • Proxy / Router & Serverless: Cohere, Groq (very high inference speed)
  • Aggregator: OpenRouter — one key, many providers
  • Each provider requires its own API key configured on the server
Read the Cloud LLM Docs
totalapp.app/settings/agentic
Cloud LLM (Hosted API)
Anthropic — Direct
Claude Sonnet 4.6, Opus 4.7, Haiku 4.5
DeepSeek — Direct
DeepSeek-V3, DeepSeek-R1
Groq — Serverless
Llama 3.3 70B at record inference speed

Run your AI where it makes sense.

Local for privacy, in-browser for zero setup, or cloud for the strongest models — TotalApp's Agentic settings let you choose, and change your mind anytime.

Get Started Free Read the Docs