> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tinfoil.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Chat models

> Chat models available on Tinfoil, using the OpenAI chat completions API.

Chat models use the OpenAI chat completions API.

<Card>
  <div style={{ display: 'flex', alignItems: 'center', gap: '16px', marginBottom: '16px' }}>
    <img src="https://mintcdn.com/tinfoil/GXlDUK5K0yNgB7O7/images/model-icons/zai.png?fit=max&auto=format&n=GXlDUK5K0yNgB7O7&q=85&s=36bb4bfa914e58deb8fbde1eba1af343" alt="Z.AI" style={{ height: '40px', width: '40px' }} width="225" height="225" data-path="images/model-icons/zai.png" />

    <div style={{ display: 'flex', alignItems: 'center', gap: '12px' }}>
      <div id="glm-5-2" className="model-anchor" style={{ fontSize: '18px', fontWeight: 'bold' }}>GLM-5.2</div>

      <a href="#glm-5-2" className="model-id-link" style={{ textDecoration: 'none' }}>
        <code className="text-xs text-gray-500 bg-gray-100 dark:bg-gray-800 dark:text-gray-300 px-1.5 py-0.5 rounded">glm-5-2</code>
      </a>
    </div>
  </div>

  **Parameters:** 754B (40B active)

  **Quantization:** FP8 (official FP8 release from Z.AI)

  **Context:** 384K tokens

  **Strengths:** State-of-the-art agentic engineering, long-horizon tool use, sustained reasoning over hundreds of iterations

  **Structured Outputs:** Structured response formatting support

  **Best for:** Agentic engineering tasks, complex coding workflows, repo-level code generation, and long-running tool-use sessions

  **Model weights:** [zai-org/GLM-5.2-FP8](https://huggingface.co/zai-org/GLM-5.2-FP8)

  **Configuration repo:** [tinfoilsh/confidential-glm5-2](https://github.com/tinfoilsh/confidential-glm5-2)
</Card>

<Card>
  <div style={{ display: 'flex', alignItems: 'center', gap: '16px', marginBottom: '16px' }}>
    <img src="https://mintcdn.com/tinfoil/HITkMM0WLVDu_kgw/images/model-icons/moonshot-light.png?fit=max&auto=format&n=HITkMM0WLVDu_kgw&q=85&s=9d0be4c66b6b9e09ab6914b854cbda2d" alt="Moonshot" style={{ height: '40px', width: '40px', flexShrink: 0 }} className="dark:block hidden" width="196" height="196" data-path="images/model-icons/moonshot-light.png" />

    <img src="https://mintcdn.com/tinfoil/HITkMM0WLVDu_kgw/images/model-icons/moonshot-dark.png?fit=max&auto=format&n=HITkMM0WLVDu_kgw&q=85&s=c5deba0135dc742cf0eec46195087a10" alt="Moonshot" style={{ height: '40px', width: '40px', flexShrink: 0 }} className="dark:hidden block" width="196" height="196" data-path="images/model-icons/moonshot-dark.png" />

    <div style={{ display: 'flex', alignItems: 'center', gap: '12px' }}>
      <div id="kimi-k2-6" className="model-anchor" style={{ fontSize: '18px', fontWeight: 'bold' }}>Kimi K2.6</div>

      <a href="#kimi-k2-6" className="model-id-link" style={{ textDecoration: 'none' }}>
        <code className="text-xs text-gray-500 bg-gray-100 dark:bg-gray-800 dark:text-gray-300 px-1.5 py-0.5 rounded">kimi-k2-6</code>
      </a>
    </div>
  </div>

  **Parameters:** 1T total (32B activated)

  **Quantization:** INT4 weight-only (native quantization-aware release from Moonshot AI)

  **Context:** 256K tokens

  **Strengths:** Long-horizon coding, image and video understanding, generates code and interfaces from visual inputs, large-scale agent orchestration, strong tool calling

  **Structured Outputs:** Structured response formatting support

  **Best for:** Agentic coding, design-to-code workflows, multimodal applications, and long-running tool-based tasks that benefit from strong reasoning

  **Model weights:** [moonshotai/Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6)

  **Configuration repo:** [tinfoilsh/confidential-kimi-k2-6](https://github.com/tinfoilsh/confidential-kimi-k2-6)

  <Info>
    **Vision + Language:** Supports text, image, and video inputs with native reasoning and tool calling for agentic workflows.
  </Info>
</Card>

<Card>
  <div style={{ display: 'flex', alignItems: 'center', gap: '16px', marginBottom: '16px' }}>
    <img src="https://mintcdn.com/tinfoil/vfMAM73wcpT2SrpR/images/model-icons/gemma.png?fit=max&auto=format&n=vfMAM73wcpT2SrpR&q=85&s=7b20f1c96812bd1c9fd512c4fead47c7" alt="Google DeepMind" style={{ height: '40px', width: '40px' }} width="225" height="225" data-path="images/model-icons/gemma.png" />

    <div style={{ display: 'flex', alignItems: 'center', gap: '12px' }}>
      <div id="gemma4-31b" className="model-anchor" style={{ fontSize: '18px', fontWeight: 'bold' }}>Gemma 4 31B</div>

      <a href="#gemma4-31b" className="model-id-link" style={{ textDecoration: 'none' }}>
        <code className="text-xs text-gray-500 bg-gray-100 dark:bg-gray-800 dark:text-gray-300 px-1.5 py-0.5 rounded">gemma4-31b</code>
      </a>
    </div>
  </div>

  **Parameters:** 31B

  **Quantization:** None (served in BF16)

  **Context:** 256K tokens

  **Strengths:** Built-in thinking mode, image understanding, native function calling, multilingual support for 35+ languages

  **Structured Outputs:** Structured response formatting support

  **Best for:** Reasoning tasks, coding, image analysis, and agentic workflows with tool calling

  **Model weights:** [google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it)

  **Configuration repo:** [tinfoilsh/confidential-gemma4-31b](https://github.com/tinfoilsh/confidential-gemma4-31b)

  <Info>
    **Vision + Language:** Processes text and image inputs. Features step-by-step reasoning with configurable thinking mode.
  </Info>
</Card>

<Card>
  <div style={{ display: 'flex', alignItems: 'center', gap: '16px', marginBottom: '16px' }}>
    <img src="https://mintcdn.com/tinfoil/7kpqELCdP4WIVCil/images/model-icons/openai-light.png?fit=max&auto=format&n=7kpqELCdP4WIVCil&q=85&s=9310e77996ef1b70d6fa1ccc4d278581" alt="OpenAI" style={{ height: '40px', width: '40px', flexShrink: 0 }} className="dark:block hidden" width="225" height="225" data-path="images/model-icons/openai-light.png" />

    <img src="https://mintcdn.com/tinfoil/7kpqELCdP4WIVCil/images/model-icons/openai.png?fit=max&auto=format&n=7kpqELCdP4WIVCil&q=85&s=1b9cbe9cd3cbfe37d282867e0d448c69" alt="OpenAI" style={{ height: '40px', width: '40px', flexShrink: 0 }} className="dark:hidden block" width="225" height="225" data-path="images/model-icons/openai.png" />

    <div style={{ display: 'flex', alignItems: 'center', gap: '12px' }}>
      <div id="gpt-oss-120b" className="model-anchor" style={{ fontSize: '18px', fontWeight: 'bold' }}>GPT-OSS 120B</div>

      <a href="#gpt-oss-120b" className="model-id-link" style={{ textDecoration: 'none' }}>
        <code className="text-xs text-gray-500 bg-gray-100 dark:bg-gray-800 dark:text-gray-300 px-1.5 py-0.5 rounded">gpt-oss-120b</code>
      </a>
    </div>
  </div>

  **Parameters:** 117B (5.1B active)

  **Quantization:** MXFP4 (native MXFP4 MoE weights, as released by OpenAI)

  **Context:** 131K tokens

  **Strengths:** Configurable reasoning effort levels, full chain-of-thought access, built-in capabilities including function calling, web browsing, and Python code execution

  **Structured Outputs:** Structured response formatting support

  **Best for:** Production use cases requiring configurable reasoning and tool use

  **Model weights:** [openai/gpt-oss-120b](https://huggingface.co/openai/gpt-oss-120b)

  **Configuration repo:** [tinfoilsh/confidential-gpt-oss-120b](https://github.com/tinfoilsh/confidential-gpt-oss-120b)
</Card>

<Card>
  <div style={{ display: 'flex', alignItems: 'center', gap: '16px', marginBottom: '16px' }}>
    <img src="https://mintcdn.com/tinfoil/7kpqELCdP4WIVCil/images/model-icons/llama.png?fit=max&auto=format&n=7kpqELCdP4WIVCil&q=85&s=0e9748643c3987fbf26b9b7531745bc9" alt="Llama" style={{ height: '40px', width: '40px' }} width="225" height="225" data-path="images/model-icons/llama.png" />

    <div style={{ display: 'flex', alignItems: 'center', gap: '12px' }}>
      <div id="llama3-3-70b" className="model-anchor" style={{ fontSize: '18px', fontWeight: 'bold' }}>Llama 3.3 70B</div>

      <a href="#llama3-3-70b" className="model-id-link" style={{ textDecoration: 'none' }}>
        <code className="text-xs text-gray-500 bg-gray-100 dark:bg-gray-800 dark:text-gray-300 px-1.5 py-0.5 rounded">llama3-3-70b</code>
      </a>
    </div>
  </div>

  **Parameters:** 70B

  **Quantization:** FP8 (FP8 weights with dynamic FP8 activations, quantized from the BF16 release)

  **Context:** 128K tokens

  **Strengths:** Multilingual, dialogue-optimized, function calling

  **Structured Outputs:** Structured response formatting support

  **Best for:** Conversational AI applications and complex dialogue systems

  **Model weights:** [RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic](https://huggingface.co/RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic)

  **Configuration repo:** [tinfoilsh/confidential-llama3-3-70b](https://github.com/tinfoilsh/confidential-llama3-3-70b)
</Card>
