Parameters: 2.8T total (104B activated)Quantization: MXFP4 weights / MXFP8 activations (native quantization-aware release from Moonshot AI)Context: 256K tokensStrengths: Long-horizon coding, complex reasoning, multimodal understanding, and agentic workflows with tool callingStructured Outputs: Structured response formatting supportBest for: Complex reasoning, repo-level coding, visual knowledge work, and long-running agentic tasksModel weights: moonshotai/Kimi-K3Configuration repo: tinfoilsh/confidential-kimi-k3
Vision + Language: Supports text and image inputs with native reasoning and tool calling for agentic workflows.

GLM-5.3
glm-5-3Reasoning is always on. Control it with
reasoning_effort (low, high, or max; defaults to max). GLM-5.3 does not honour enable_thinking; use reasoning_effort: "low" for the lightest reasoning. Reasoning tokens count toward max_tokens, so structured outputs and strict tool calls at the default effort need a generous max_tokens (around 20K) or a lower effort.
GLM-5.3 Flash
glm-5-3-flashVision and reasoning: Supports text and image inputs; see the Image Processing Guide. Reasoning is always on: control it with
reasoning_effort (low, high, or max; defaults to max). GLM-5.3 Flash does not honour enable_thinking; use reasoning_effort: "low" for the lightest reasoning. Reasoning tokens count toward max_tokens, so structured outputs and strict tool calls at the default effort need a generous max_tokens (around 20K) or a lower effort.
DeepSeek V4 Flash
deepseek-v4-flash
DeepSeek V4.1 Flash
deepseek-v4-1-flashlow, high, xhigh, max), built-in speculative decoding for fast generation, tool callingStructured Outputs: Structured response formatting supportBest for: Long-document and image-grounded analysis, large-codebase understanding, and high-volume agentic workloads over very long inputsModel weights: deepseek-ai/DeepSeek-V4.1-FlashConfiguration repo: tinfoilsh/confidential-deepseek-v4-1-flash
Gemma 4 31B
gemma4-31bVision + Language: Processes text and image inputs. Features step-by-step reasoning with configurable thinking mode.
Parameters: 117B (5.1B active)Quantization: MXFP4 (native MXFP4 MoE weights, as released by OpenAI)Context: 131K tokensStrengths: Configurable reasoning effort levels, full chain-of-thought access, built-in capabilities including function calling, web browsing, and Python code executionStructured Outputs: Structured response formatting supportBest for: Production use cases requiring configurable reasoning and tool useModel weights: openai/gpt-oss-120bConfiguration repo: tinfoilsh/confidential-gpt-oss-120b

Llama 3.3 70B
llama3-3-70b




