Multi-Provider AI Workspace & Local AI Gateway
Maximize your AI productivity with Local AI through Local Device GGUF (100% offline & private on your PC) or your own Local Cloud Modal GPU endpoint, or connect your API keys for OpenAI, Anthropic, DeepSeek, Qwen, Kimi AI, Cloudflare Workers AI, Agnes AI, Google Gemini & Groq LPU with our 12-Tier Auto RollOver Router. Use one local AI platform for chat, image generation, and coding workflows across compatible clients.
Next-Generation Multi-Engine Features
Explore a local-first architecture that combines Local Device GGUF, an optional user-owned Local Cloud Modal endpoint, Direct BYOK providers, and a 12-Tier cloud fallback cascade.
Local AI — Device or Cloud (Priority 1)
Choose Local Device to run GGUF models 100% offline with CUDA, Vulkan, or CPU, or select Local Cloud to connect HyperRouter to your own Modal-managed GPU inference endpoint.
Smart Hardware Acceleration
Automatically selects available CUDA, Vulkan, or CPU runtime assets for local GGUF execution.
Premium Image Studio
Generate stunning visuals with multiple aspect ratios (1:1, 16:9, 9:16, 4:3). Features automatic failover between Google Cloud and Pollinations AI networks.
Antigravity Code Workspace
An autonomous multi-turn agentic terminal powered by remote Linux Cloud Sandboxes. Perfect for complex software engineering and algorithm design.
Intelligent Budgeting & Probes
Automated real-time billing probes detect Free vs Paid Tier accounts with Eco, Maximum, and Auto prompt allocation strategies to control costs.
Reasoning UI Trace Parser
Real-time Gemma 4 & Groq token evaluator intercepting reasoning streams inside SSE streams, displayed in an interactive glassmorphic accordion.
How AlphaOne AI Works
Bring Your Own Key (BYOK) keeps provider credentials and quota choices under your control.
AUTHORIZED AI ACCESS ONLY
Official APIs, user-owned credentials and provider-permitted local models. No scraping, no access-control bypassing, and no unauthorized wrappers or free-access workarounds.
Choose Your AI Source
Local Device: Point HyperRouter to a folder containing .gguf model files for offline execution with no cloud API key required.
Local Cloud: Connect your own Modal-managed inference endpoint and credentials when you want remote GPU compute.
Cloud AI: Obtain an API key or token from supported official providers such as Google AI Studio, OpenAI, Anthropic, DeepSeek, Qwen, Kimi AI, Cloudflare Workers AI, Agnes AI, or GroqCloud. Free and paid quotas depend on each provider and your account.
Connect Securely
Enter supported provider credentials in Settings. HyperRouter stores its Windows gateway configuration in the current Windows user profile, while cloud requests are sent only to the provider selected by the routing engine.
Route Across Local and Cloud AI
Start chatting, creating artwork, or writing code. HyperRouter evaluates the selected Local AI source when enabled, then routes eligible failures through Direct Providers and the 12-Tier fallback cascade.
Supported AI Models & Failover Hierarchy
| Provider | Tier | Model ID / Identifier | Description & Capabilities |
|---|---|---|---|
| Google Direct | Direct Tier 0 | gemini-3.8-flash-high / medium / low | Gemini 3.8 Flash with dashboard-controlled High, Medium, or Low thinking level for Paid accounts. |
| OpenAI Direct | Direct Tier 0 | gpt-5.6-sol / terra / luna | Direct OpenAI flagship models with instant failover priority. |
| Anthropic Direct | Direct Tier 0 | claude-opus-5 / claude-sonnet-5 / claude-haiku-4-5-20251001 | Direct Anthropic Claude models with instant failover priority. |
| DeepSeek Direct | Direct Tier 0 | deepseek-v4-pro / flash | Direct DeepSeek high-throughput reasoning & code synthesis models. |
| Kimi AI Direct | Direct Tier 0 | kimi-k3 / kimi-k2.7-code / highspeed | Direct Kimi AI (Moonshot) 1M context deep reasoning & specialized coding models. |
| Qwen Direct | Direct Tier 0 | qwen3.8-max / qwen3.8-flash / qwen3.7-plus | Direct Qwen Dashscope reasoning & coding models with custom Base URL. |
| Custom Direct | Direct Tier 0 | User-configured Model ID Example: gemma-4-26b-a4b |
Direct custom OpenAI-compatible endpoint with custom Base URL, API Key, and Model ID configured in HyperRouter Settings. |
| Cloudflare Workers AI Direct | Direct Tier 0 | User-selected Workers AI Model ID Example: @cf/qwen/qwen3.8-27b |
Direct Cloudflare Workers AI with a free-form Model ID selected by the user in HyperRouter Settings. The same selected model is used for High, Medium, Low, and Custom profiles, and the provider participates in Direct Provider Auto RollOver. |
| Tier 1 | gemini-3.7-flash | Flagship primary model supporting up to 1M context architecture (Token limits governed by Google API quota). | |
| Tier 2 | gemini-3.6-flash | High-speed, high-capacity primary fallback model (1M Context). | |
| Tier 3 | gemma-4-31b-it | Massive open-weights reasoning model from Google. | |
| Tier 4 | gemma-4-26b-a4b-it | Specialized open-weights coding model from Google. | |
| Tier 5 | gemini-3.5-flash | High-efficiency general reasoning fallback model. | |
| Tier 6 | gemini-3.5-flash-lite | High-speed lightweight Gemini engine for rapid response. | |
| Tier 7 | gemini-3-flash-preview | Experimental preview model for advanced syntax analysis. | |
| Agnes AI | Tier 8 | agnes-2.5-flash | Ultra-large 512K context window & coding specialist buffer. |
| Groq LPU | Tier 9 | llama-3.3-70b-versatile | Ultra-fast 70B parameter open model on Groq LPU hardware. |
| Groq LPU | Tier 10 | llama-3.1-8b-instant | High-throughput instant-response fallback layer on Groq LPU. |
| Groq LPU | Tier 11 | qwen/qwen3.6-27b | High-speed coding and multilingual specialist. |
| Groq LPU | Tier 12 | openai/gpt-oss-20b | High-efficiency open reasoning engine for low-latency streaming. |
🛠️ Modern Technology Stack
Built on .NET 10.0 Blazor WebAssembly and C# 13 with strict Clean Architecture, SSE streaming, and FontAwesome vector assets. The web workspace and the Windows HyperRouter gateway remain separate execution surfaces.