> For the complete documentation index, see [llms.txt](https://docs.blockbrain.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.blockbrain.ai/for-users/all-about-llms/overview-of-llms.md).

# Overview of LLMs

Explore and compare the most popular Large Language Models (LLMs) from GPT to Claude and beyond.

### 1. Find Primary Use Cases LLMs & R-LLMs

<table><thead><tr><th width="282.18182373046875">Use Case</th><th>LLM Models</th></tr></thead><tbody><tr><td>General Productivity</td><td>GPT 5.6 Terra, Gemini 3.5 Flash, Claude Sonnet 5</td></tr><tr><td>Complex Reasoning</td><td>Claude Opus 5, GPT 5.6 Sol, Gemini 3.5 Flash</td></tr><tr><td>Structured Writing &#x26; Synthesis</td><td>Claude Sonnet 5, GPT 5.6 Terra, Gemini 3.1 Pro</td></tr><tr><td>Coding &#x26; Technical Workflows</td><td>Claude Opus 4.8, Claude Sonnet 5, GPT 5.3 Codex</td></tr><tr><td>Fast &#x26; Scalable Processing</td><td>GPT 5.6 Luna, Gemini 3.5 Flash (Lite), Claude Haiku 4.5</td></tr></tbody></table>

### 2. Hosting Preference

Choose where your data is processed based on your privacy needs and access priorities. We offer two hosting options, EU and US, each with different benefits around compliance, speed, and model access.

| Factor                       | EU Hosting (Privacy First)                                  | US Hosting (Feature First)                          |
| ---------------------------- | ----------------------------------------------------------- | --------------------------------------------------- |
| **GDPR Compliance**          | Fully GDPR-compliant                                        | Not GDPR-compliant by default                       |
| **Data Residency**           | Data stays in the EU                                        | Data stored globally                                |
| **Model Availability**       | Late model release depending on EU Data Center availability | Full access to the latest models and features first |
| **Legal & Regulatory Risks** | Meets stricter EU privacy laws                              | Subject to US law and transfer safeguards           |

**Summary:**

* **Choose EU hosting** if you prioritize **GDPR compliance** and **strict data privacy** within Europe.
* **Choose US hosting** if you want **the latest models** and global data centers.

### 3. Speed vs. Depth: What Matters More to You?

Some models are designed for quick, lightweight tasks. Others are built to dive deeper, think harder, and handle more complexity. Choose based on the kind of experience you need.

| Preference                                  | When to Choose                                                                     | Models                                                   |
| ------------------------------------------- | ---------------------------------------------------------------------------------- | -------------------------------------------------------- |
| <p>High Speed<br>(Fast, Responsive)</p>     | For fast answers or simple tasks where low latency matters most.                   | Gemini 3.5 Flash, Claude Sonnet 4.6 (Fast), GPT 5.6 Luna |
| <p>High Depth<br>(Detailed, Structured)</p> | For complex prompts, multi-step logic, or detailed analysis that needs reflection. | Gemini 3.5 Flash, Claude Opus 5, GPT 5.6 Sol             |

### 4. Choose your Desired LLMs

### AWS Bedrock: Anthropic Claude

| Model                    | Description                                                                                                                                                                                                                                                                                                                                                                                  |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Claude Opus 5            | A major step up from Opus 4.8 at the same price, coming close to Fable 5 intelligence at half the cost. It more than doubles Opus 4.8 on Frontier-Bench v0.1 (43.3% vs 21.1%) and leads on knowledge work with 1861 on GDPval-AA v2. Best for long-running agents and complex professional work where the model needs to verify its own output.                                              |
| Claude Sonnet 5          | The most agentic Sonnet thus-far and a substantial upgrade over Sonnet 4.6 across reasoning, tool use, coding, and knowledge work. It comes close to Opus 4.8 at a lower price, with adjustable effort levels that can match Opus on tasks like agentic search and computer use. Best for multi-step coding and automation work needing near-flagship capability without flagship cost.      |
| Claude Opus 4.8          | A step up from Opus 4.7 across the board. It is a top model for agentic coding and computer use. It leads on SWE-Bench Pro (69.2%) against GPT-5.5 and Gemini 3.1 Pro, and scores 83.4% on OSWorld-Verified for agentic computer use. A key highlight is its improved honesty, making it a more reliable partner for complex, long-running tasks.                                            |
| Claude Opus 4.7          | It is a major step up from Opus 4.6, built for difficult hands-on coding work. It scores 64.5% on SWE-bench Pro (up from Opus 4.6's 53.4%), and users report being able to hand off their hardest engineering tasks with confidence. Beyond coding, it's a reliable upgrade across all metrics.                                                                                              |
| Claude Opus 4.6          | A step up from Opus 4.5 on coding, with more reliable performance on agentic tasks, codebase management, and structured code review and debugging. It also handles heavy work better, staying more consistent and efficient across analysis, document review, and multitasking.                                                                                                              |
| Claude Opus 4.6 (Max)    | A configuration of Claude Opus 4.6. It operates in "Max Reasoning" mode. It applies full reasoning depth for the most complex tasks, maximizing accuracy and reliability with higher latency and cost.                                                                                                                                                                                       |
| Claude Opus 4.6 (High)   | A configuration of Claude Opus 4.6. It operates in "High Reasoning" mode. It uses deeper reasoning steps for more complex tasks, improving accuracy and structure at the cost of speed.                                                                                                                                                                                                      |
| Claude Opus 4.6 (Medium) | A configuration of Claude Opus 4.6. It operates in "Medium Reasoning" mode. It balances reasoning depth with speed, allowing for more structured outputs without significantly increasing latency.                                                                                                                                                                                           |
| Claude Opus 4.6 (Low)    | A configuration of Claude Opus 4.6. It operates in "Low Reasoning" mode. It minimizes extended reasoning steps for faster, more efficient outputs.                                                                                                                                                                                                                                           |
| Claude Sonnet 4.6        | Great everyday model, stronger than Sonnet 4.5 at coding, computer use, and long-context reasoning, scoring 79.6% on SWE-bench Verified. Well suited for agent workflows, large-codebase tasks, document analysis, and other multi-step work.                                                                                                                                                |
| Claude Sonnet 4.6 (Fast) | A configuration of Claude Sonnet 4.6. It operates strictly in "Non-Reasoning" mode, bypassing long thinking steps in order to have low-latency quality outputs.                                                                                                                                                                                                                              |
| Claude Haiku 4.5         | Fast, efficient, and built for scale. Delivers near-Sonnet-level coding and reasoning at about 3× cheaper and 2× faster performance. Excels in tool use, UI interaction, and parallel task execution, ideal as a worker model in multi-agent or production setups. Strong on coding reliability. Best for backend automations, chat workloads, and agent systems needing speed and low cost. |
| Claude Haiku 4.5 (Fast)  | Fast, efficient, and built for scale. Delivers near-Sonnet-level coding and reasoning at about 3× cheaper and 2× faster performance. Excels in tool use, UI interaction, and parallel task execution, ideal as a worker model in multi-agent or production setups. Strong on coding reliability. Best for backend automations, chat workloads, and agent systems needing speed and low cost. |

### AWS Bedrock: Other

| Model       | Description                                                                                                                                                                                                                                                                                                                                                          |
| ----------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Gemma 4 31B | A powerful open-weight model built for advanced reasoning, coding, and agentic workflows. The 31B Dense variant ranks #3 among open models on Arena AI, outperforming some models up to 20× larger. Supports native image and video understanding, and 140+ languages. Best for self-hosted coding assistants, document analysis, and customizable enterprise agents |

### Google Vertex AI: Gemini

| Model                              | Description                                                                                                                                                                                                                                                                                                                                                                                                                             |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Gemini 3.5 Flash                   | A major step up from Gemini 3.1 Pro. This is built for agentic and coding work at Flash speed. It leads on MCP Atlas (83.6%) and Finance Agent v2 (57.9%), outperforming GPT-5.5 and Claude Opus 4.7 on those benchmarks, while also topping multimodal understanding at 84.2% on CharXiv Reasoning. At 4x the output speed of other frontier models, it's the go-to choice for powerful and fast multi-step task handling.             |
| Gemini 3.5 Flash (High)            | A configuration of Gemini 3.5 Flash with reasoning effort set to high. Best for difficult coding, agentic workflows and multi-step reasoning.                                                                                                                                                                                                                                                                                           |
| Gemini 3.5 Flash (Medium)          | A configuration of Gemini 3.5 Flash with reasoning effort set to medium. It balances thinking depth and response speed for reliable everyday performance. Best as the general default when tasks vary in complexity.                                                                                                                                                                                                                    |
| Gemini 3.5 Flash (Low)             | A configuration of Gemini 3.5 Flash with reasoning effort set to low. It answers quickly with light reasoning, keeping cost and latency down.                                                                                                                                                                                                                                                                                           |
| Gemini 3.5 Flash (Minimal)         | A configuration of Gemini 3.5 Flash with reasoning effort set to minimal. Best for classification, extraction and quick responses.                                                                                                                                                                                                                                                                                                      |
| Gemini 3.1 Pro                     | Great for complex reasoning, coding, and long-context work across text, images, audio, video, PDFs, and large codebases. Best used for deep research, multi-step agent workflows, technical planning, and document-heavy analysis where strong reasoning and broad multimodal understanding matter most.                                                                                                                                |
| Gemini 3.1 Flash (Lite)            | The fastest and most cost-efficient model in the Gemini 3 series. Very helpful for high-volume tasks. It outperforms comparable models like GPT-5 Mini and Claude 4.5 Haiku across reasoning and multimodal benchmarks including 86.9% on GPQA Diamond and 76.8% on MMMU-Pro. Best used for translation, content moderation, data labeling, and any task you need to run at scale quickly and affordably.                               |
| Gemini 3 Pro                       | High-performance multimodal model designed for teams that work with text, images, documents, videos and code. It offers strong long-context reasoning, reliable analysis across large files and advanced tool use for more automated workflows. For businesses that rely on visual data, technical documentation or development tasks, this provides broader multimodal coverage and deeper file understanding than text-focused models |
| Gemini 3 Flash                     | (no description in JSON)                                                                                                                                                                                                                                                                                                                                                                                                                |
| Gemini 2.5 Pro                     | Well-rounded, reliable model for everyday tasks. Excels at reasoning, conversation, and coding, making it a strong fit for smart assistants, business tools, and creative workflows that require speed, accuracy and thoughtful output.                                                                                                                                                                                                 |
| Gemini 2.5 Pro (Enterprise Search) | Excels at reasoning, conversation, and coding, and is configured for secure commercial use. It uses Web Grounding for Enterprise to provide access to live web data without the privacy risks of standard search tools.                                                                                                                                                                                                                 |
| Gemini 2.5 Flash                   | Smarter and more capable than 2.0 Flash, with fast responses and visible reasoning. Best for real-time tasks that need both speed and lightweight thinking, like chatbots, copilots, and scalable AI tools.                                                                                                                                                                                                                             |
| Gemini Live 2.5 Flash Native Audio | Real-time voice conversations with native audio input and output. Ideal for voice assistants and live user interactions, with a focus on speed and cost over deep reasoning.                                                                                                                                                                                                                                                            |
| Gemini 2.5 Flash (Lite)            | Gemini 2.5 Flash-Lite delivers fast, affordable, real-time performance—ideal for high-volume, everyday tasks like chatbots, customer support, and quick content creation. It’s built for speed and efficiency, offering smarter answers without the heavy overhead of deep, complex reasoning.                                                                                                                                          |
| Gemini 2.0 Flash                   | Built for fast, responsive performance even with large inputs. Ideal for real-time apps, smart assistants, and systems that need instant answers across long conversations or documents. Supports up to 1 million tokens, making it a strong choice for tasks that require both speed and scale.                                                                                                                                        |
| Gemini 2.0 Flash (Thinking Mode)   | Designed for academic precision, complex logic, and structured problem-solving. Well-suited for scientific reports, step-by-step reasoning, and data interpretation tasks that require depth over speed. Not intended for fast-turnaround or real-time use like the standard Flash model.                                                                                                                                               |
| Gemini 2.0 Flash (Lite)            | A streamlined, cost-efficient version of Gemini Flash, built for fast, high-volume interactions. Best for chatbots, summaries, and real-time tasks where speed and scale matter more than deep reasoning or complex context handling.                                                                                                                                                                                                   |
| Gemini 1.5 Pro                     | Impressive 1M token context window. Balanced model with very good price-performance ratio. Strong multimodal capabilities.                                                                                                                                                                                                                                                                                                              |
| (no name in JSON)                  | Predecessor to Gemini 1.5. Solid basic performance for various tasks. Good multimodal capabilities.                                                                                                                                                                                                                                                                                                                                     |

### OpenAI

| Model                               | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| GPT 5.6 Sol                         | OpenAI's flagship and best model to date, a step up from GPT-5.5 across coding, knowledge work, and computer use. It scores 80 on the Artificial Analysis Coding Agent Index and 64.6% on SWE-Bench Pro (up from 59.4%), while working faster and using fewer tokens. Best for difficult, long-running tasks that need high quality outputs.                                                                                                                                                                                   |
| GPT 5.6 Terra                       | The balanced middle tier of the GPT-5.6 family, matching or beating GPT-5.5 at half the price. It competes with GPT-5.5 on coding with 77.4 on the Artificial Analysis Coding Agent Index (vs 76.4). The practical default for everyday work needing near-flagship quality without more efficient cost.                                                                                                                                                                                                                        |
| GPT 5.6 Luna                        | The fastest and cheapest model in the GPT-5.6 family, built for speed and volume. It nearly matches GPT-5.5's peak performance at less than half the cost, and beats Claude Opus 4.8 on the Artificial Analysis Coding Agent Index. Best for high-volume, cost-sensitive tasks that still need reliable quality.                                                                                                                                                                                                               |
| GPT 5.5 Pro                         | Uses the same underlying model as GPT-5.5, but uses parallel test-time compute to push further on the hardest tasks. Leads on BrowseComp (90.1%) and FrontierMath Tier 4 (39.6%), making it the strongest option for deep research and advanced mathematics. Requests take longer to complete and are priced higher. Reserve this for problems where standard reasoning effort is not enough.                                                                                                                                  |
| GPT 5.5                             | Handles complex agentic work well alongside coding, research, computer use and high token tasks. Scores 82.7% on Terminal-Bench 2.0 and 84.9% on GDPval, leading all competing models at GPT-5.4 per-token latency. Best used for tasks that requires sustained reasoning across context.                                                                                                                                                                                                                                      |
| GPT 5.5 Instant                     | An upgrade over GPT-5.3 Instant. It is built to be faster and more affordable than the full GPT-5.5 model. It significantly cuts down on hallucinations, with inaccuracies reduced by 37.3% on real user-flagged conversations. Useful for GPT-5.5-level intelligence for everyday tasks with cheaper costs over the full model.                                                                                                                                                                                               |
| GPT 5.4                             | Handles complex, reasoning-heavy work really well. More reliable and consistent than GPT-5.3, especially for multi-step tasks, structured outputs, and unclear prompts.                                                                                                                                                                                                                                                                                                                                                        |
| GPT 5.4 (High Thinking with Search) | Uses more reasoning resources for complex and multi-step tasks. Best for deep analysis, strategy, and technical problem-solving. Supports automatic search for up-to-date information, with higher latency and cost.                                                                                                                                                                                                                                                                                                           |
| GPT 5.4 (High Thinking)             | A configuration of GPT 5.4. It operates in "High-Thinking. It uses extended thinking steps for more refined results.                                                                                                                                                                                                                                                                                                                                                                                                           |
| GPT 5.4 (Low Thinking with Search)  | Uses additional reasoning compute to improve accuracy and structure. Suitable for summarization, light analysis, structured outputs, and basic decision-making. Can automatically search for up-to-date information when needed.                                                                                                                                                                                                                                                                                               |
| GPT 5.4 (Low Thinking)              | A configuration of GPT 5.4. It operates in "Low-Thinking" mode. It skips extended thinking steps for faster results.                                                                                                                                                                                                                                                                                                                                                                                                           |
| GPT 5.4 (Non-Thinking with Search)  | Uses baseline reasoning without extra compute. Optimized for speed and cost, ideal for simple tasks like formatting, classification, and structured outputs, with automatic search when needed.                                                                                                                                                                                                                                                                                                                                |
| GPT 5.4 Mini                        | Mainly designed for coding, tool workflows, and high-volume execution tasks. This model is built for handling debugging, codebase navigation, and fast iteration loops. It is fast and cost efficient. Best used as a default production model rather than for deep reasoning.                                                                                                                                                                                                                                                 |
| GPT 5.4 Nano                        | The latest small and fast model in the GPT-5.4 family, built for tasks where speed and cost matter more than deep reasoning. It's an upgrade over GPT-5 Nano, and excels at high-volume work like classification, data extraction, and acting as a supporting agent within larger AI workflows. If you're running tasks at scale and need quick, reliable outputs without burning through credits, Nano is the right pick.                                                                                                     |
| GPT 5.3 Codex                       | Built specifically for agentic software engineering, combining the coding performance of GPT-5.2 Codex with broader reasoning across research, documentation, and technical decision-making, running 25% faster than its predecessor. Best used for long-running coding tasks, debugging, and complex execution workflows.                                                                                                                                                                                                     |
| GPT 5.2                             | Built for end-to-end professional knowledge work and long-running agents, with stronger reliability across complex workflows. It produces higher-quality work artifacts (spreadsheets, slides), writes and refactors code more effectively, understands long context better, perceives images more accurately, and uses tools to complete multi-step projects. Designed as the most capable upgrade in the GPT-5 series, GPT-5.2 is optimized for turning instructions into finished outputs that deliver real economic value. |
| GPT 5.2 (High Thinking)             | A configuration of GPT-5.2. It operates in "High Thinking" mode. It uses deeper reasoning steps for more complex tasks, improving accuracy at the cost of speed.                                                                                                                                                                                                                                                                                                                                                               |
| GPT 5.2 (Low Thinking with Search)  | A configuration of GPT-5.2. It operates in "Low Thinking" mode with search enabled. It limits extended reasoning while using retrieval to support fast, up-to-date responses.                                                                                                                                                                                                                                                                                                                                                  |
| GPT 5.2 (Low Thinking)              | A configuration of GPT-5.2. It operates in "Low Thinking" mode. It skips extended thinking steps for faster, more efficient outputs.                                                                                                                                                                                                                                                                                                                                                                                           |
| GPT 5.2 (Non-Thinking with Search)  | A configuration of GPT-5.2. It operates in "No Thinking" mode with search enabled. It skips reasoning steps and relies on retrieval for fast, up-to-date responses.                                                                                                                                                                                                                                                                                                                                                            |
| GPT 5.2 (Thinking with Search)      | A configuration of GPT-5.2. It operates in "Thinking" mode with search enabled. It uses deeper reasoning while leveraging retrieval to improve accuracy and keep responses up to date.                                                                                                                                                                                                                                                                                                                                         |
| GPT 5                               | Tailored for natural, enterprise-grade conversations. Multimodal and context-aware, it flows seamlessly across long dialogues, handles images and text, and adapts tone and reasoning depth. Note: Intended to replace GPT 4o and GPT 4.1                                                                                                                                                                                                                                                                                      |
| GPT 5 (Thinking)                    | Thinks like an expert, adapts to the task. Handles everything from writing and coding to research and data analysis, shifting seamlessly between quick answers and in-depth reasoning. Works with text and images, keeps context over long sessions, and delivers more accurate, reliable results. Note: Intended to replace OpenAI o3 and OpenAI o3 Pro                                                                                                                                                                       |
| GPT 5 Mini                          | Lean, cost-conscious, and still sharp. Delivers reliable instruction-following, rich multimodal responses (text + images), and reduced latency. All while sharing GPT‑5’s strong reasoning and safety tuning, but at a lighter compute and price point. Note: Intended to replace GPT 4o Mini and OpenAI o4 Mini                                                                                                                                                                                                               |
| GPT 5 Nano                          | Ultra-light and lightning‑fast. Built for low-latency needs like summarization, classification, and quick Q\&A, it taps GPT‑5’s reasoning smarts at a fraction of the cost and compute. Note: Intended to replace GPT 4.1 Nano                                                                                                                                                                                                                                                                                                 |
| GPT 4.5                             | Caution: Very expensive — 10–15× more costly than GPT-4o. Delivers subtle improvements in emotional tone, writing flow, and creative ideation, especially in chat-like settings. Best for polished conversation output, but not recommended for reasoning, analysis, or complex tasks.                                                                                                                                                                                                                                         |
| (no name in JSON)                   | Specialized version of GPT-4 with outstanding capabilities in image analysis and understanding. Can explain and analyze complex visual concepts.                                                                                                                                                                                                                                                                                                                                                                               |
| (no name in JSON)                   | Optimized GPT-4 version with higher processing speed at slightly reduced precision. More current training knowledge than base GPT-4.                                                                                                                                                                                                                                                                                                                                                                                           |
| GPT 4 Omni                          | Designed for real-time interaction across text, audio, and images. Ideal for smart assistants, live applications, and creative tasks that benefit from fast responses and seamless multimodal input. Offers faster performance and lower cost than earlier GPT-4 models.                                                                                                                                                                                                                                                       |
| OpenAI GPT 4o Realtime Audio        | Ideal for fast, conversational tasks with streaming audio. Extremely low latency and strong comprehension for voice-based interactions. Best for real-time assistants and audio-first products.                                                                                                                                                                                                                                                                                                                                |
| GPT 4o Mini                         | A smaller, more affordable version of GPT-4 Omni, designed to give fast, high-quality responses without the higher cost. Great for smart assistants, everyday tasks, and apps that need reliable reasoning with some image or audio input.                                                                                                                                                                                                                                                                                     |
| GPT Realtime 2                      | The newest live voice model. Useful for real conversations that require actual reasoning. It's a significant upgrade over GPT-Realtime-1.5, jumping from 81.4% to 96.6% on Big Bench Audio, and is best suited for voice agents.                                                                                                                                                                                                                                                                                               |
| OpenAI GPT Realtime                 | Built for live, low-latency conversations with voice-in, voice-out streaming. It uses GPT-4o’s reasoning core but is optimized for fast dialogue. Best for interactive assistants, real-time translators, and conversational AI with natural audio responses. Strong in speed and usability, but not intended for long, complex reasoning tasks.                                                                                                                                                                               |
| o4 Mini                             | A faster, lower-cost version of o4 built for quick conversations and real-time responsiveness. Ideal for chatbots, simple assistants, and everyday tools that need speed, clarity, and a light touch without the resource demands of larger models.                                                                                                                                                                                                                                                                            |
| o3 Pro                              | DISCLAIMER: Very Expensive! Use only if necessary. Long waiting times depending on the prompt, making it unsuitable for real-time use. Great for high-stakes legal analysis, deep business strategy, and technical decision-making. Efficient at breaking down complex documents, building logic-driven reports, and guiding long-term plans.                                                                                                                                                                                  |
| o3                                  | CAREFUL: Very cost-intensive. Great for logic-heavy tasks like legal analysis, business strategy, and technical planning. It breaks down complex documents, guides long-term plans, and creates structured, insight-driven outputs.                                                                                                                                                                                                                                                                                            |
| o3 Mini                             | A smaller version of o3, optimized for fast, structured responses. Ideal for short-form reasoning, technical prompts, and lightweight tasks where speed and efficiency matter more than creative depth.                                                                                                                                                                                                                                                                                                                        |
| o1                                  | Warning: Very expensive! Specialized for smart problem-solving. Great for technical reasoning, coding help, and structured thinking tasks, but less conversational for casual users.                                                                                                                                                                                                                                                                                                                                           |
| o1 Mini                             | Warning: Very Expensive! A compact reasoning model. It is best for basic technical tasks, faster lightweight deployments, and logic workflows that don't need deep explanations.                                                                                                                                                                                                                                                                                                                                               |

### Azure OpenAI

| Model                       | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| GPT 5.6 Sol                 | OpenAI's flagship and best model to date, a step up from GPT-5.5 across coding, knowledge work, and computer use. It scores 80 on the Artificial Analysis Coding Agent Index and 64.6% on SWE-Bench Pro (up from 59.4%), while working faster and using fewer tokens. Best for difficult, long-running tasks that need high quality outputs.                                                                                                                                                                                   |
| GPT 5.6 Terra               | The balanced middle tier of the GPT-5.6 family, matching or beating GPT-5.5 at half the price. It competes with GPT-5.5 on coding with 77.4 on the Artificial Analysis Coding Agent Index (vs 76.4). The practical default for everyday work needing near-flagship quality without more efficient cost.                                                                                                                                                                                                                        |
| GPT 5.6 Luna                | The fastest and cheapest model in the GPT-5.6 family, built for speed and volume. It nearly matches GPT-5.5's peak performance at less than half the cost, and beats Claude Opus 4.8 on the Artificial Analysis Coding Agent Index. Best for high-volume, cost-sensitive tasks that still need reliable quality.                                                                                                                                                                                                               |
| GPT 5.5 Instant             | An upgrade over GPT-5.3 Instant. It is built to be faster and more affordable than the full GPT-5.5 model. It significantly cuts down on hallucinations, with inaccuracies reduced by 37.3% on real user-flagged conversations. Useful for GPT-5.5-level intelligence for everyday tasks with cheaper costs over the full model.                                                                                                                                                                                               |
| GPT 5.4                     | Handles complex, reasoning-heavy work really well. More reliable and consistent than GPT-5.3, especially for multi-step tasks, structured outputs, and unclear prompts.                                                                                                                                                                                                                                                                                                                                                        |
| GPT 5.4 (High Thinking)     | A configuration of GPT 5.4. It operates in "High-Thinking. It uses extended thinking steps for more refined results.                                                                                                                                                                                                                                                                                                                                                                                                           |
| GPT 5.4 (Low Thinking)      | Configuration of GPT 5.4 that runs in a low-thinking mode. It skips extended reasoning steps to return faster responses, optimized for low latency while maintaining solid output quality for straightforward tasks.                                                                                                                                                                                                                                                                                                           |
| GPT 5.2                     | Built for end-to-end professional knowledge work and long-running agents, with stronger reliability across complex workflows. It produces higher-quality work artifacts (spreadsheets, slides), writes and refactors code more effectively, understands long context better, perceives images more accurately, and uses tools to complete multi-step projects. Designed as the most capable upgrade in the GPT-5 series, GPT-5.2 is optimized for turning instructions into finished outputs that deliver real economic value. |
| GPT 5.1                     | Built for fast, reliable conversations with stronger instruction-following and more natural dialogue. It responds quickly to simple tasks, maintains clarity in longer threads, and adapts tone with more control. Designed as a refined upgrade to GPT-5 with smoother communication and improved everyday usability.                                                                                                                                                                                                         |
| GPT 5.1 Codex               | High value for developer workflows like code generation and programming tasks. Best for short-to-medium coding problems that require clear instruction following. Works especially well for human-in-the-loop development.                                                                                                                                                                                                                                                                                                     |
| GPT 5.1 Codex Max           | Built for complex, high-risk coding work and autonomous agents. Stronger reasoning and consistency across large codebases than standard Codex. Suitable for multi-step refactors and long-context work, with higher latency and cost. Best used selectively for engineering tasks where correctness and depth outweigh speed.                                                                                                                                                                                                  |
| GPT 5.1 (High Thinking)     | Uses high “reasoning” token resources for deeper answers. Best for work that requires planning, judgment, and careful thinking. Includes the full knowledge base and core capabilities of GPT-5.1.                                                                                                                                                                                                                                                                                                                             |
| GPT 5.1 (Low Thinking)      | Utilizes reduced 'reasoning' resources for high responsiveness and efficiency. It bypasses deep reflective processing to deliver near-instant answers, making it ideal for real-time conversation and high-volume data processing. It prioritizes speed for workflows where complex problem-solving is not required.                                                                                                                                                                                                           |
| GPT 5.1 (Non-Thinking)      | This utilizes the least "reasoning" resources in order to provide the lowest possibly latency and maximum cost efficiency. It is designed for tasks where analysis is less necessary such as simple text formatting, high-volume classification, and conversational flow while still leveraging the full knowledge base and core capabilities of GPT 5.1.                                                                                                                                                                      |
| GPT 5                       | Tailored for natural, enterprise-grade conversations. Multimodal and context-aware, it flows seamlessly across long dialogues, handles images and text, and adapts tone and reasoning depth. Note: Intended to replace GPT 4o and GPT 4.1                                                                                                                                                                                                                                                                                      |
| GPT 5 Codex                 | GPT-5 Codex is a version of OpenAI’s GPT-5 model tailored for software engineering and coding tasks. It serves as an AI coding assistant that can build projects, add features, debug, refactor code, and perform code reviews. The model adapts its reasoning time to the complexity of tasks, and supports multimodal inputs like images.                                                                                                                                                                                    |
| GPT 5 (Thinking)            | Thinks like an expert, adapts to the task. Handles everything from writing and coding to research and data analysis, shifting seamlessly between quick answers and in-depth reasoning. Works with text and images, keeps context over long sessions, and delivers more accurate, reliable results. Note: Successor to o3 and o3 Pro                                                                                                                                                                                            |
| GPT 5 Mini                  | Lean, cost-conscious, and still sharp. Delivers reliable instruction-following, rich multimodal responses (text + images), and reduced latency. All while sharing GPT‑5’s strong reasoning and safety tuning, but at a lighter compute and price point. Note: Intended to replace GPT 4o Mini and OpenAI o4 Mini                                                                                                                                                                                                               |
| GPT 5 Nano                  | Ultra-light and lightning‑fast. Built for low-latency needs like summarization, classification, and quick Q\&A, it taps GPT‑5’s reasoning smarts at a fraction of the cost and compute. Note: Intended to replace GPT 4.1 Nano                                                                                                                                                                                                                                                                                                 |
| GPT 5 (Auto)                | Automatically selects the most suitable model from the GPT model family based on your query's complexity and performance needs. Saves you time from manually switching between different models while optimizing for quality, speed, and efficiency. The selected model is shown with each response. Routes to: GPT 5 & GPT 4.1 family and o4 Mini. **This Model is currently in Preview by Azure AI.**                                                                                                                        |
| GPT 4.1                     | A powerful and refined model designed for tasks that demand precision and depth. This excels in coding, instruction-following, and handling extensive documents, thanks to its expanded 1 million-token context window. It's a reliable choice for developers, researchers, and professionals seeking enhanced performance and cost-efficiency over earlier GPT-4 models.                                                                                                                                                      |
| GPT 4.1 Mini                | A faster, more efficient version of GPT-4.1 that keeps much of its quality while using fewer resources. Great for developers, startups, product teams, and anyone building smart tools or assistants that need quick, capable performance at lower cost and with faster response times.                                                                                                                                                                                                                                        |
| GPT 4.1 Nano                | The fastest and most lightweight GPT-4.1 model, built for speed and simplicity. Ideal for developers, product teams, and startups needing quick autocomplete, fast classifications, or lightweight assistants where cost and response time matter most.                                                                                                                                                                                                                                                                        |
| Azure GPT 4o Realtime Audio | Ideal for fast, conversational tasks with streaming audio. Extremely low latency and strong comprehension for voice-based interactions. Best for real-time assistants and audio-first products.                                                                                                                                                                                                                                                                                                                                |
| GPT 4 Omni                  | Designed for real-time interaction across text, audio, and images. Ideal for smart assistants, live applications, and creative tasks that benefit from fast responses and seamless multimodal input. Offers faster performance and lower cost than earlier GPT-4 models.                                                                                                                                                                                                                                                       |
| GPT 4 Turbo                 | Optimized GPT-4 version with higher processing speed at slightly reduced precision. More current training knowledge than base GPT-4.                                                                                                                                                                                                                                                                                                                                                                                           |
| GPT 4 Vision                | Specialized version of GPT-4 with outstanding capabilities in image analysis and understanding. Can explain and analyze complex visual concepts.                                                                                                                                                                                                                                                                                                                                                                               |
| GPT 4o Mini                 | A smaller, more affordable version of GPT-4 Omni, designed to give fast, high-quality responses without the higher cost. Great for smart assistants, everyday tasks, and apps that need reliable reasoning with some image or audio input.                                                                                                                                                                                                                                                                                     |
| GPT 3.5 Turbo               | Proven model for standard tasks. Very fast and efficient. Good for simple to medium requirements.                                                                                                                                                                                                                                                                                                                                                                                                                              |
| (no name in JSON)           | (no description in JSON)                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| DeepSeek R1                 | Warning: Developed in China and Hosted in the US. Open-source model built for deep reasoning, logical problem solving, and structured analysis. Highly effective in math and scientific tasks, and ideal for explainable, customizable applications.                                                                                                                                                                                                                                                                           |
| Azure GPT Realtime          | Built for live, low-latency conversations with voice-in, voice-out streaming. It uses GPT-4o’s reasoning core but is optimized for fast dialogue. Best for interactive assistants, real-time translators, and conversational AI with natural audio responses. Strong in speed and usability, but not intended for long, complex reasoning tasks.                                                                                                                                                                               |
| o4 Mini                     | A faster, lower-cost version of o4 built for quick conversations and real-time responsiveness. Ideal for chatbots, simple assistants, and everyday tools that need speed, clarity, and a light touch without the resource demands of larger models.                                                                                                                                                                                                                                                                            |
| o3 Pro                      | DISCLAIMER: Very Expensive! Use only if necessary. Long waiting times depending on the prompt, making it unsuitable for real-time use. Great for high-stakes legal analysis, deep business strategy, and technical decision-making. Efficient at breaking down complex documents, building logic-driven reports, and guiding long-term plans.                                                                                                                                                                                  |
| o3                          | CAREFUL: Very cost-intensive. Great for logic-heavy tasks like legal analysis, business strategy, and technical planning. It breaks down complex documents, guides long-term plans, and creates structured, insight-driven outputs.                                                                                                                                                                                                                                                                                            |
| o3 Mini                     | A smaller version of o3, optimized for fast, structured responses. Ideal for short-form reasoning, technical prompts, and lightweight tasks where speed and efficiency matter more than creative depth.                                                                                                                                                                                                                                                                                                                        |
| o1                          | Warning: Very expensive! Specialized for smart problem-solving. Great for technical reasoning, coding help, and structured thinking tasks, but less conversational for casual users.                                                                                                                                                                                                                                                                                                                                           |
| o1 Mini                     | Warning: Very Expensive! A compact reasoning model. It is best for basic technical tasks, faster lightweight deployments, and logic workflows that don't need deep explanations.                                                                                                                                                                                                                                                                                                                                               |

### Google Vertex AI: Anthropic Claude

| Model                        | Description                                                                                                                                                                                                                                                                                                                                                                                  |
| ---------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Claude Opus 5                | A major step up from Opus 4.8 at the same price, coming close to Fable 5 intelligence at half the cost. It more than doubles Opus 4.8 on Frontier-Bench v0.1 (43.3% vs 21.1%) and leads on knowledge work with 1861 on GDPval-AA v2. Best for long-running agents and complex professional work where the model needs to verify its own output.                                              |
| Claude Sonnet 5              | The most agentic Sonnet thus-far and a substantial upgrade over Sonnet 4.6 across reasoning, tool use, coding, and knowledge work. It comes close to Opus 4.8 at a lower price, with adjustable effort levels that can match Opus on tasks like agentic search and computer use. Best for multi-step coding and automation work needing near-flagship capability without flagship cost.      |
| Claude Opus 4.8              | A step up from Opus 4.7 across the board, its the new top model for agentic coding and computer use. It leads on SWE-Bench Pro (69.2%) against GPT-5.5 and Gemini 3.1 Pro, and scores 83.4% on OSWorld-Verified for agentic computer use. A key highlight is its improved honesty, making it a more reliable partner for complex, long-running tasks.                                        |
| Claude Opus 4.7              | This model is currently in public preview, meaning it's available for early access in EU. It is a major step up from Opus 4.6, built for difficult hands-on coding work. It scores 64.5% on SWE-bench Pro (up from Opus 4.6's 53.4%), and users report being able to hand off their hardest engineering tasks with confidence. Beyond coding, it's a reliable upgrade across all metrics.    |
| Claude Opus 4.6              | Improves on the coding strengths of prior models with more reliable performance for agentic tasks, codebase management, and structured code review and debugging. It’s also stronger in everyday work. From handling analysis, document review, and multitasking more consistently and efficiently.                                                                                          |
| Claude Opus 4.6 (Max)        | A configuration of Claude Opus 4.7. It operates in "Max Reasoning" mode. It applies full reasoning depth for the most complex tasks, maximizing accuracy and reliability with higher latency and cost.                                                                                                                                                                                       |
| Claude Opus 4.6 (High)       | A configuration of Claude Opus 4.7. It operates in "High Reasoning" mode. It uses deeper reasoning steps for more complex tasks, improving accuracy and structure at the cost of speed.                                                                                                                                                                                                      |
| Claude Opus 4.6 (Medium)     | A configuration of Claude Opus 4.7. It operates in "Medium Reasoning" mode. It balances reasoning depth with speed, allowing for more structured outputs without significantly increasing latency.                                                                                                                                                                                           |
| Claude Opus 4.6 (Low)        | A configuration of Claude Opus 4.7. It operates in "Low Reasoning" mode. It minimizes extended reasoning steps for faster, more efficient outputs.                                                                                                                                                                                                                                           |
| Claude Sonnet 4.6            | Improved Sonnet model with stronger coding, computer use, and long-context reasoning. Better suited for agent workflows, large-codebase tasks, document analysis, and other multi-step work.                                                                                                                                                                                                 |
| Claude Sonnet 4.6 (Fast)     | A configuration of Claude Sonnet 4.6. It operates strictly in "Non-Reasoning" mode, bypassing long thinking steps in order to have low-latency quality outputs.                                                                                                                                                                                                                              |
| Claude Opus 4.5              | Built for agentic coding, long-horizon reasoning, and complex tool-driven workflows, with high correctness across multi-step execution. It excels at software engineering tasks, terminal and computer use, structured tool orchestration, and novel problem solving, achieving strong first-pass success with fewer retries.                                                                |
| Claude Sonnet 4.5            | Extends balanced performance with stronger reasoning, longer task persistence, and more reliable tool use, achieving state-of-the-art results in coding and real-world computer tasks. Delivers higher accuracy, greater autonomy, and smoother multi-step workflows compared to Claude 4 Sonnet.                                                                                            |
| Claude Sonnet 4.5 (Fast)     | A configuration of Claude Sonnet 4.5. It operates strictly in "Non-Reasoning" mode, bypassing long thinking steps in order to have low-latency quality outputs.                                                                                                                                                                                                                              |
| Claude Haiku 4.5             | Fast, efficient, and built for scale. Delivers near-Sonnet-level coding and reasoning at about 3× cheaper and 2× faster performance. Excels in tool use, UI interaction, and parallel task execution, ideal as a worker model in multi-agent or production setups. Strong on coding reliability. Best for backend automations, chat workloads, and agent systems needing speed and low cost. |
| Claude Opus 4.1              | Sharper, steadier, smarter. Builds on Opus 4 with noticeably cleaner code fixes, longer autonomous agentic workflows, and more precise reasoning. Tackles long, multi-step tasks with hybrid reasoning and extended "thinking", including code refactoring, deep research, and strategic synthesis.                                                                                          |
| Claude Opus 4                | Designed for long, demanding tasks that require deep thinking and consistent performance. It’s especially strong at powering autonomous agents, handling large, complex codebases, and running multi-step workflows that can last for hours. It is ideal for advanced coding tasks, detailed analysis, and tasks that need smart, sustained focus.                                           |
| Claude Sonnet 4              | Enhances balanced performance through refined reasoning, extended autonomy, and precise instruction-following, delivering improved coding capabilities and greater task accuracy compared to Claude 3.7 Sonnet.                                                                                                                                                                              |
| Claude 3.7 Sonnet            | Designed for complex document analysis, coding tasks, and multi-step reasoning. Well-suited for problem-solving agents and analysis bots that handle long-form inputs and technical workflows. Often excessive for basic or lightweight queries.                                                                                                                                             |
| Claude 3.7 Sonnet (Thinking) | Designed for deep, deliberate reasoning and critical thinking. Excels in slow, multi-step analysis for strategic tasks like legal review, policy generation, and complex customer support.                                                                                                                                                                                                   |
| Claude 3.5 Sonnet            | Focused on structured writing, instruction following, and business-grade accuracy. Well-suited for report generation, formal communication, and smart document workflows. Efficient for medium-length content where clarity and consistency matter.                                                                                                                                          |
| Claude 3.5 Sonnet v2         | Optimized for automating business processes and structured workflows. A solid choice for enterprise tools, DevOps support, and productivity tasks where speed, clarity, and consistent output are key.                                                                                                                                                                                       |
| Claude 3.5 Haiku             | Optimized for fast, high-volume tasks. This is great for quick replies, customer chatbots, and real-time document summaries, but less suited for deep reasoning or complex creative work.                                                                                                                                                                                                    |
| Claude 3 Opus                | Powerful legacy model for logic-heavy work. Ideal for enterprises, researchers, and anyone working with complex text and image data.                                                                                                                                                                                                                                                         |
| Claude 3 Sonnet              | Balanced version with good ratio between speed and accuracy. Sufficient for most business applications.                                                                                                                                                                                                                                                                                      |
| Claude 3 Haiku               | Efficient version of Claude 3, optimized for fast processing. Maintains the core capabilities of the Claude 3 family with reduced complexity. Ideal for real-time applications and quick analyses.                                                                                                                                                                                           |

### Anthropic

| Model                             | Description                                                                                                                                                                                                                                                                                                                                                                                  |
| --------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Claude Opus 4.7                   | (no description in JSON)                                                                                                                                                                                                                                                                                                                                                                     |
| Claude Opus 4.6                   | Improves on the coding strengths of prior models with more reliable performance for agentic tasks, codebase management, and structured code review and debugging. It’s also stronger in everyday work. From handling analysis, document review, and multitasking more consistently and efficiently.                                                                                          |
| Claude Sonnet 4.6                 | Improved Sonnet model with stronger coding, computer use, and long-context reasoning. Better suited for agent workflows, large-codebase tasks, document analysis, and other multi-step work.                                                                                                                                                                                                 |
| Claude Sonnet 4.6 (Fast)          | A configuration of Claude Sonnet 4.6. It operates strictly in "Non-Reasoning" mode, bypassing long thinking steps in order to have low-latency quality outputs.                                                                                                                                                                                                                              |
| Claude Opus 4.5                   | Built for agentic coding, long-horizon reasoning, and complex tool-driven workflows, with high correctness across multi-step execution. It excels at software engineering tasks, terminal and computer use, structured tool orchestration, and novel problem solving, achieving strong first-pass success with fewer retries.                                                                |
| Claude Sonnet 4.5                 | Extends balanced performance with stronger reasoning, longer task persistence, and more reliable tool use, achieving state-of-the-art results in coding and real-world computer tasks. Delivers higher accuracy, greater autonomy, and smoother multi-step workflows compared to Claude 4 Sonnet.                                                                                            |
| Claude Sonnet 4.5 (Fast)          | A configuration of Claude Sonnet 4.5. It operates strictly in "Non-Reasoning" mode, bypassing long thinking steps in order to have low-latency quality outputs.                                                                                                                                                                                                                              |
| Claude Haiku 4.5                  | Fast, efficient, and built for scale. Delivers near-Sonnet-level coding and reasoning at about 3× cheaper and 2× faster performance. Excels in tool use, UI interaction, and parallel task execution, ideal as a worker model in multi-agent or production setups. Strong on coding reliability. Best for backend automations, chat workloads, and agent systems needing speed and low cost. |
| Claude Opus 4.1                   | Sharper, steadier, smarter. Builds on Opus 4 with noticeably cleaner code fixes, longer autonomous agentic workflows, and more precise reasoning. Tackles long, multi-step tasks with hybrid reasoning and extended "thinking", including code refactoring, deep research, and strategic synthesis.                                                                                          |
| Claude Opus 4                     | Designed for long, demanding tasks that require deep thinking and consistent performance. It’s especially strong at powering autonomous agents, handling large, complex codebases, and running multi-step workflows that can last for hours. It is ideal for advanced coding tasks, detailed analysis, and tasks that need smart, sustained focus.                                           |
| Claude Sonnet 4                   | Enhances balanced performance through refined reasoning, extended autonomy, and precise instruction-following, delivering improved coding capabilities and greater task accuracy compared to Claude 3.7 Sonnet.                                                                                                                                                                              |
| Claude 3.7 Sonnet                 | Designed for complex document analysis, coding tasks, and multi-step reasoning. Well-suited for problem-solving agents and analysis bots that handle long-form inputs and technical workflows. Often excessive for basic or lightweight queries.                                                                                                                                             |
| Claude 3.7 Sonnet (Thinking Mode) | Designed for deep, deliberate reasoning and critical thinking. Excels in slow, multi-step analysis for strategic tasks like legal review, policy generation, and complex customer support.                                                                                                                                                                                                   |
| Claude 3.5 Sonnet                 | Focused on structured writing, instruction following, and business-grade accuracy. Well-suited for report generation, formal communication, and smart document workflows. Efficient for medium-length content where clarity and consistency matter.                                                                                                                                          |
| Claude 3.5 Sonnet v2              | Optimized for automating business processes and structured workflows. A solid choice for enterprise tools, DevOps support, and productivity tasks where speed, clarity, and consistent output are key.                                                                                                                                                                                       |
| Claude 3 Opus                     | Powerful legacy model for logic-heavy work. Ideal for enterprises, researchers, and anyone working with complex text and image data.                                                                                                                                                                                                                                                         |
| Claude 3 Sonnet                   | Balanced version with good ratio between speed and accuracy. Sufficient for most business applications.                                                                                                                                                                                                                                                                                      |
| Claude 3 Haiku                    | Efficient version of Claude 3, optimized for fast processing. Maintains the core capabilities of the Claude 3 family with reduced complexity. Ideal for real-time applications and quick analyses.                                                                                                                                                                                           |

### Mistral

| Model              | Description                                                                                                                                                                                                                                                                                                         |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Mistral Small 4    | (no description in JSON)                                                                                                                                                                                                                                                                                            |
| Mistral Medium 3.5 | Mistral's open-weight model, merging old reasoning and coding lines into one set of weights. Strong at agentic coding and multi-tool work, scoring 77.6% on SWE-bench Verified.                                                                                                                                     |
| Mistral Large 3    | Mistral's generalist model. This is a non-reasoning model built for high-volume straightforward work rather than hard reasoning. Cheap and widely hosted, but it sits well below the reasoning tier on quality benchmarks.                                                                                          |
| Magistral Medium   | An enterprise-grade reasoning powerhouse. Built from the ground up via reinforcement learning, it delivers fast, traceable, multilingual chains-of-thought—think “flash answers” that explain how they think. Handles complex tasks in math, code, logic, and rule-based workflows with precision and transparency. |
| Mistral Large      | A multilingual powerhouse built for deep reasoning, code, math, and agent workflows. Offers a 32K‑token context window, native function‑calling, and RAG support. Thinks precisely across multiple languages, and delivers high performance on benchmarks.                                                          |
| Mistral Medium     | A frontier-class generalist built for enterprise demands. Excels in programming, mathematical reasoning, long-document comprehension, dialogue, and multimodal tasks. Offers extended context (up to 128K tokens), function calling, and agentic workflows.                                                         |
| Mistral Small      | Lean, swift, and powerful. Built for rapid, high-volume tasks with low latency. Features strong instruction-following, conversational finesse, and in version 3.1, robust multimodal understanding.                                                                                                                 |

### Google Vertex AI: Mistral

| Model               | Description                                                                                                                                                                                                                                                                               |
| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| AI21 Jamba Large    | Great for legal analysis, enterprise research, and multilingual document tasks. Handles long-form content with strong memory, improved reasoning, and broad language support which is ideal for deep analysis across global teams.                                                        |
| AI21 Jamba Mini     | Compact model for simple tasks. Very fast with limited complexity. Good for basic applications.                                                                                                                                                                                           |
| Mistral Small 3.1   | Lightweight, open-source multimodal model built for fast and cost-efficient deployment. Handles long-context workloads up to 128K tokens, and function calling. Well-suited for on-device use cases, responsive virtual assistants, document and image processing, and agentic workflows. |
| Mistral Medium 3    | Mistral's generalist model. This is a non-reasoning model built for high-volume straightforward work rather than hard reasoning. Cheap and widely hosted. It scores 19 on the Artificial Analysis Intelligence Index, slightly average for non-reasoning models in its price tier         |
| Mistral Codestral 2 | Specialized code generation model, built for precise completion and fill-in-the-middle tasks. It supports writing, editing, and conversing about code across many programming languages.                                                                                                  |
| Mistral Codestral   | Well-suited for real-time development, code automation, and debugging tasks across various programming languages. Improves code generation speed, accuracy, and multi-language support over earlier open models like Code LLaMA, StarCoder, and DeepSeek-Coder.                           |
| Mistral Large       | Efficient model with strong reasoning. Particularly good with technical and scientific texts. Cost-effective alternative to top models.                                                                                                                                                   |
| Mistral Nemo        | Focuses on language fluency, response quality, and multilingual coverage, offering stronger performance than open-source chat models like LLaMA 2 Chat. Well-suited for chatbots, writing assistance, and global customer support.                                                        |

### Google Vertex AI: Llama

| Model                | Description                                                                                                                                                                                                                                                                                           |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Llama 4 Maverick 17B | Balanced power for demanding, mixed workloads. Combines strong multimodal reasoning, coding, and multilingual skills with a 1-million token context. Best for: versatile enterprise apps, advanced chat systems, and AI agents handling both technical and creative tasks.                            |
| Llama 4 Scout 17B    | Designed for long-form analysis, deep document comprehension, and multimodal tasks. Supports an immense 10-million token context and excels in reasoning, image understanding, and coding. Best for large-document workflows, visual reasoning, and teams seeking high capability on modest hardware. |
| Llama 3.2 90B        | Understands both language and images with impressive accuracy. Ideal for research, enterprise tools, and tasks that combine visual reasoning with strong language skills.                                                                                                                             |
| Llama 3.1 405B       | Largest Llama model. Strong in general text processing. Good performance in technical tasks. Open-source base.                                                                                                                                                                                        |
| Llama 3.1 70B        | Efficient Llama model for general applications. Good balance between size and performance.                                                                                                                                                                                                            |
| Llama 3.1 8B         | Smallest Llama model with reduced context window. Suitable for simple text processing. Very efficient with limited resources.                                                                                                                                                                         |

### xAI Grok

| Model                                | Description                                                                                                                                                                                                                                                                                                                                                                              |
| ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Grok 4.5                             | xAI’s latest flagship model for coding, agentic tasks, and knowledge work. Strong on real-world engineering, scoring 83.3% on Terminal-Bench 2.1 and 64.7% on SWE-Bench Pro, while offering fast-model speeds and competitive pricing. Best for coding agents, technical workflows, app building, and business productivity tasks that need strong reasoning with lower latency and cost |
| Grok 4                               | Frontier-level Grok model with native tool use, real-time web access, and advanced reasoning. Ideal for power users who need live data, long-context comprehension, and complex task execution. Great for research and code to strategic decision-making. Built to compete with GPT-4o and Gemini 1.5 Pro at the highest tier.                                                           |
| Grok 3                               | Built for stronger reasoning and deeper conversations across everyday and technical topics. Well-suited for smart assistants, research tools, and advanced chat experiences that aim to compete with top-tier models like GPT-4 and Claude Opus.                                                                                                                                         |
| Grok 3 Mini (Thinking - High Effort) | A more efficient Grok model tuned for thoughtful, deeper reasoning. Ideal for moderately complex tasks where you want accurate, well-considered answers without the full cost of a flagship model. Leading to a smart balance between speed, accuracy, and effort.                                                                                                                       |
| Grok 3 Mini (Thinking - Low Effort)  | A fast, low-resource Grok model designed for casual interactions and simple tasks. Ideal when speed matters more than depth. This is perfect for everyday chat, lightweight assistants, and high-volume use where quick answers are the priority.                                                                                                                                        |
| Grok 2 Vision                        | Designed to understand both images and text, making it helpful for visual queries, screenshots, and image-based tasks. Great for users who need smart answers grounded in what they see.                                                                                                                                                                                                 |

### DeepSeek

| Model              | Description                                                                                                                                                                                                                                |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| DeepSeek Chat (V3) | Warning: Hosted and developed in China. Fast, open-source assistant built for everyday use. Well-suited for customer-facing apps, developer tools, and plug-and-play AI that handles smart conversations and basic coding tasks.           |
| Reasoner (R1)      | Warning: Hosted and developed in China. Open-source model built for deep reasoning, logical problem solving, and structured analysis. Highly effective in math and scientific tasks, and ideal for explainable, customizable applications. |

### Nebius

| Model                  | Description                                                                                                                                                                                                       |
| ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| DeepSeek Chat (V3)     | Chinese Model: Fast, open-source assistant built for everyday use. Well-suited for customer-facing apps, developer tools, and plug-and-play AI that handles smart conversations and basic coding tasks.           |
| DeepSeek Reasoner (R1) | Chinese Model: Open-source model built for deep reasoning, logical problem solving, and structured analysis. Highly effective in math and scientific tasks, and ideal for explainable, customizable applications. |

### STACKIT

| Model         | Description                                                                                                                                                                                                                                                                                                               |
| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| GPT-OSS 120B  | OpenAI’s largest open-weight reasoning model, built for strong agentic work. It matches or exceeds o4-mini across coding, reasoning, tool calling, competition and math. Strong fit for enterprise agents, coding workflows, and specialized use cases that require data control.                                         |
| GPT-OSS 20B   | A compact open-weight reasoning model built for low-latency local and specialized deployments. Matches or exceeds o3-mini on common coding, math, and tool-use evaluations, while running with as little as 16 GB of memory. Strong fit for private agent workflows, on-device inference, and cost-conscious deployments. |
| Qwen3.6 27B   | An open-weight multimodal model that performs well above its size in agentic coding. It scores 77.2% on SWE-bench Verified, 53.5% on SWE-bench Pro, and 59.3% on Terminal-Bench 2.0, making it a strong option for coding, frontend workflows, tool use, and visual tasks.                                                |
| Llama 3.3 70B | Balanced, high-performance model that delivers robust reasoning and reliable output quality. It performs well in complex tasks such as analysis, long-form content creation, and contextual decision support.                                                                                                             |
| Gemma 3 27B   | A capable open-weight multimodal model. Supports text and image inputs, a 128K context window, 140+ languages, function calling, and structured outputs. Strong fit for document analysis, multilingual assistants, visual tasks, and flexible agentic workflows.                                                         |
| Qwen3 VL 235B | A powerful open-weight vision-language model built for advanced multimodal work. Its native 256K context makes it well suited for long documents, hours-long video, complex visual analysis, and agentic workflows. Best for teams that need high-end multimodal capabilities and can support a heavyweight deployment    |

### IONOS

| Model             | Description                                                                                                                                                                                                                                                                                                                 |
| ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| GPT-OSS 120B      | This is an EU-sovereign transformer model designed to deliver strong reasoning, contextual understanding, and high-quality generative output. As a 120B-parameter system running entirely on IONOS infrastructure, it offers a balance of power and performance while meeting strict European data protection requirements. |
| Mistral Small 24B | A strong blend of reasoning quality and performance in a mid-size architecture. It handles complex instructions, multi-step tasks, and general conversational workloads with high reliability.                                                                                                                              |
| Mistral Nemo 12B  | Compact and capable model optimized for efficiency and high-throughput performance. Built on the NeMo and Mistral open-source stack, it excels in fast, lightweight tasks such as summarization, classification, and structured or template-driven generation.                                                              |
| Teuken 7B         | A lightweight, multilingual model optimized for efficiency and cost-effective usage. It supports basic conversational tasks, simple content generation, etc.                                                                                                                                                                |
| Llama 3.3 70B     | Balanced, high-performance model that delivers robust reasoning and reliable output quality. It performs well in complex tasks such as analysis, long-form content creation, and contextual decision support.                                                                                                               |
| Llama 3.1 405B    | High-capacity model from Meta’s LLaMA 3.1 family, designed for scenarios requiring maximum expressiveness and broad linguistic capability. Its large parameter count gives it strong generalization, nuanced generation, and broad language coverage.                                                                       |
| Llama 3.1 8B      | Provides fast, efficient inference suited for everyday operational tasks. It handles straightforward instructions, short-form content, and utility workflows with high responsiveness. The model’s low resource footprint makes it ideal for cost-optimized deployments.                                                    |

### Pharia / Aleph Alpha

| Model               | Description                                                                                                                                                                                                                                                                                                  |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| GPT-OSS 120B        | Provided via Aleph Alpha. This is an open-weight reasoning model by OpenAI that achieves near-parity with OpenAI o4-mini on core reasoning benchmarks. Well-suited for production-grade tasks where transparency and sovereign hosting matter.                                                               |
| Kimi K2.5           | Provided via Aleph Alpha. This is an open-source native multimodal model by Moonshot AI that excels at coding, visual reasoning, and multi-step agentic workflows. It is ideal for autonomous research and document workflows.                                                                               |
| Aleph Alpha Preview | Sovereign European model, trained and served entirely on EU infrastructure, currently in public preview. It targets German-language and regulated workloads where data residency drives the decision (with a 33k context window). Best for German-language tasks under strict data sovereignty requirements. |


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.blockbrain.ai/for-users/all-about-llms/overview-of-llms.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
