Skip to main content
Chat models use the OpenAI chat completions API.
MoonshotMoonshot
Kimi K3
kimi-k3
Parameters: 2.8T total (104B activated)Quantization: MXFP4 weights / MXFP8 activations (native quantization-aware release from Moonshot AI)Context: 256K tokensStrengths: Long-horizon coding, complex reasoning, multimodal understanding, and agentic workflows with tool callingStructured Outputs: Structured response formatting supportBest for: Complex reasoning, repo-level coding, visual knowledge work, and long-running agentic tasksModel weights: moonshotai/Kimi-K3Configuration repo: tinfoilsh/confidential-kimi-k3
Vision + Language: Supports text and image inputs with native reasoning and tool calling for agentic workflows.
Z.AI
GLM-5.3
glm-5-3
Parameters: 743B (39B active)Quantization: NVFP4 MoE experts (Inferact release; attention, shared experts, and dense layers kept in BF16)Context: 1M tokensStrengths: Frontier agentic coding and long-horizon tool use, with always-on reasoning and configurable effortStructured Outputs: Structured response formatting supportBest for: Agentic engineering, repo-scale code generation and refactoring, and long-running tool-use sessions over very long contextModel weights: Inferact/GLM-5.3-NVFP4Configuration repo: tinfoilsh/confidential-glm5-3-nvfp4
Reasoning is always on. Control it with reasoning_effort (low, high, or max; defaults to max). GLM-5.3 does not honour enable_thinking; use reasoning_effort: "low" for the lightest reasoning. Reasoning tokens count toward max_tokens, so structured outputs and strict tool calls at the default effort need a generous max_tokens (around 20K) or a lower effort.
Z.AI
GLM-5.3 Flash
glm-5-3-flash
Parameters: 320B (18B active)Quantization: NVFP4 transformer linear layers (Red Hat AI release; vision tower, embeddings, and output head kept in original precision)Context: 1M tokensStrengths: Fast generation with built-in speculative decoding, very long context, image input, always-on reasoning with configurable effort, tool callingStructured Outputs: Structured response formatting supportBest for: Long-document and multimodal tasks, agentic tool use, and high-volume workloads where speed and cost matterModel weights: RedHatAI/GLM-5.3-Flash-NVFP4Configuration repo: tinfoilsh/confidential-glm5-3-flash
Vision and reasoning: Supports text and image inputs; see the Image Processing Guide. Reasoning is always on: control it with reasoning_effort (low, high, or max; defaults to max). GLM-5.3 Flash does not honour enable_thinking; use reasoning_effort: "low" for the lightest reasoning. Reasoning tokens count toward max_tokens, so structured outputs and strict tool calls at the default effort need a generous max_tokens (around 20K) or a lower effort.
DeepSeek
DeepSeek V4 Flash
deepseek-v4-flash
Version: 0731 releaseParameters: 304B (MoE, 6 of 257 experts active per token)Quantization: FP8 attention / FP4 experts (native quantized release from DeepSeek)Context: 1M tokensStrengths: Very long context, optional reasoning with configurable effort, built-in speculative decoding for fast generation, tool callingStructured Outputs: Structured response formatting supportBest for: Long-document analysis, large-codebase understanding, and high-volume agentic workloads over very long inputsModel weights: deepseek-ai/DeepSeek-V4-Flash-0731Configuration repo: tinfoilsh/confidential-deepseek-v4-flash
DeepSeek
DeepSeek V4.1 Flash
deepseek-v4-1-flash
Experimental: this model is not yet fully productionized. Capacity and performance may change while we finish rolling it out.
Version: V4.1 Flash releaseParameters: 552B (MoE, 6 of 384 routed experts plus 1 shared active per token; Engram memory)Quantization: FP8 attention / FP4 experts (native quantized release from DeepSeek)Context: 1M tokensStrengths: Very long context, image input, optional reasoning with configurable effort (low, high, xhigh, max), built-in speculative decoding for fast generation, tool callingStructured Outputs: Structured response formatting supportBest for: Long-document and image-grounded analysis, large-codebase understanding, and high-volume agentic workloads over very long inputsModel weights: deepseek-ai/DeepSeek-V4.1-FlashConfiguration repo: tinfoilsh/confidential-deepseek-v4-1-flash
Google DeepMind
Gemma 4 31B
gemma4-31b
Parameters: 31BQuantization: None (served in BF16)Context: 256K tokensStrengths: Built-in thinking mode, image understanding, native function calling, multilingual support for 35+ languagesStructured Outputs: Structured response formatting supportBest for: Reasoning tasks, coding, image analysis, and agentic workflows with tool callingModel weights: google/gemma-4-31B-itConfiguration repo: tinfoilsh/confidential-gemma4-31b
Vision + Language: Processes text and image inputs. Features step-by-step reasoning with configurable thinking mode.
OpenAIOpenAI
GPT-OSS 120B
gpt-oss-120b
Parameters: 117B (5.1B active)Quantization: MXFP4 (native MXFP4 MoE weights, as released by OpenAI)Context: 131K tokensStrengths: Configurable reasoning effort levels, full chain-of-thought access, built-in capabilities including function calling, web browsing, and Python code executionStructured Outputs: Structured response formatting supportBest for: Production use cases requiring configurable reasoning and tool useModel weights: openai/gpt-oss-120bConfiguration repo: tinfoilsh/confidential-gpt-oss-120b
Llama
Llama 3.3 70B
llama3-3-70b
Parameters: 70BQuantization: FP8 (FP8 weights with dynamic FP8 activations, quantized from the BF16 release)Context: 128K tokensStrengths: Multilingual, dialogue-optimized, function callingStructured Outputs: Structured response formatting supportBest for: Conversational AI applications and complex dialogue systemsModel weights: RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamicConfiguration repo: tinfoilsh/confidential-llama3-3-70b