Pro Tips

The AI Model
Cheat Sheet

The 8 AI models worth knowing right now, what each is genuinely best at, the honest weakness of each, and what it costs. Plus the switch-between-models workflow that beats picking a favorite.

Everyone keeps asking me the same thing. Which AI model is the best one? I get why people want a single answer. But I think it is the wrong question to ask.

The better question is not which model is best. It is which model is best for this task. The people getting the most out of AI stopped picking a favorite. They switch between models based on the job in front of them, and once I saw the pattern I could not unsee it.

So here is the actual cheat sheet. The eight models worth knowing right now, what each one is genuinely good at, and the honest weakness of each. Save this one.

Cheat SheetThe 8 Models, At A Glance

Scroll sideways on your phone to see the whole table. Cost is per million tokens, shown as input and output, which is how these are actually billed.

ModelReach for it whenReal weaknessCost in / out
Claude Fable 5The hardest problem you have. Best coding scores of anything you can use right now.Most expensive on this list. Overkill for everyday work.$10 / $50
Claude Opus 4.8Long jobs that run for hours. Big migrations, deep research, agents that hold context. Great natural writing.Pricey, and slower than you want for quick questions.$5 / $25
Claude Sonnet 5Your daily driver. Plans, uses a browser and terminal, runs on its own.Uses more tokens per task, so at max effort it can cost more than Opus.$2 / $10
GPT-5.6 SolHard multi-step problems and coding that spans files, tests, and follow-up fixes.Priciest of the GPT tiers, and you rarely need this ceiling.$5 / $30
GPT-5.6 TerraThe broad middle of real work. Support replies, document analysis, drafting at scale.Not the top ceiling. Hard reasoning goes to Sol.$2.50 / $15
GPT-5.6 LunaSummarizing, labeling, extracting, first drafts. One fifth the cost of Sol.Long-document recall falls off a cliff. Do not hand it a huge file.$1 / $6
Gemini 3.5 FlashSpeed, and anything needing live Google Search grounding.Gemini writing can read a little flat next to Claude.$1.50 / $9
Grok 4.3Breaking news and live X, plus video you want it to actually watch.Trails the others on coding, and takes about 12 seconds to start answering.$1.25 / $2.50

Prices checked July 2026. Sonnet 5's intro pricing rises about 50% on September 1, 2026.

If You Only Remember Four Things

Hardest problem you have: Claude Fable 5.
Everyday work: Claude Sonnet 5 or GPT-5.6 Terra.
Cheap and high volume: GPT-5.6 Luna or Gemini 3.5 Flash.
What happened in the last hour: Grok 4.3.

By TaskLook Up Your Job, Get The Model

This is the one to screenshot. Find what you are actually doing, use the first pick, and fall back to the second if you do not have access to it.

What you are doingUse thisOr thisWhy
Writing people will read
blogs, newsletters, emails in your voice
Claude Opus 4.8Claude Sonnet 5Most natural prose, and it holds your tone over a long piece.
Short punchy copy
hooks, subject lines, ads, captions
GPT-5.6 TerraGPT-5.6 SolGPT is stronger at short business copy than long-form voice.
Hardest coding problemClaude Fable 5Claude Opus 4.8Top coding score of any model you can use today.
Everyday coding and code reviewClaude Sonnet 5GPT-5.6 TerraHandles most real work without the flagship price.
Agents and terminal workGPT-5.6 SolClaude Sonnet 5Sol genuinely leads on agentic and terminal operation.
A job that runs for hours
migrations, deep research
Claude Opus 4.8Claude Fable 5Built to hold context across long autonomous work.
Work you will not double-checkClaude Opus 4.8Grok 4.3Both are unusually willing to say they are unsure instead of inventing an answer.
Research that must be currentGemini 3.5 FlashGrok 4.3Most reliable live Google Search grounding.
Breaking news, live sentimentGrok 4.3Gemini 3.5 FlashReads X live rather than a stale index.
Big documents and spreadsheetsGemini 3.5 FlashClaude Opus 4.8Chews through large files fast without slowing down.
Video you want understoodGrok 4.3Gemini 3.5 FlashTakes video directly, up to five minutes, and follows what happens.
Bulk simple tasks
summaries, labeling, extracting
GPT-5.6 LunaGemini 3.5 FlashCheapest way to run volume. Just never on long documents.
The ModelsPros And Cons Of Each

The table gets you moving. This part is the why, so you can make the call yourself when a new model shows up next month.

Claude Fable 5

Anthropic · the ceiling · $10 / $50

Good At

  • The hardest coding job you have. It holds the top coding score of any model you can currently use.
  • Problems that beat everything else. When another model has already stalled twice, this is the escalation.
  • Planning and architecture where getting the shape wrong costs you weeks.

Not Good At

  • Everyday questions. It costs double Opus 4.8, which is already the expensive one.
  • High volume work. Running your bulk tasks through this will wreck your bill fast.
  • Anything you need back in seconds. It thinks before it answers.

Buy it if: you hit a wall the other models cannot get past, and being right matters more than the cost.

Claude Opus 4.8

Anthropic · the long-haul worker · $5 / $25

Good At

  • Jobs that run for hours without you. Big migrations, deep research, agents that hold context across a long task.
  • Writing people will actually read. The most natural long-form prose of anything here, and the best at holding your voice.
  • Being honest. It is the strongest of these at saying it is unsure instead of confidently making something up, which matters enormously when you cannot check the work.
  • Production quality code where correctness beats speed.

Not Good At

  • Speed. It is roughly half as fast as GPT-5.6 Sol on the same task.
  • Cheap bulk output. It produces close to three times more output per task than Sol, so output-heavy work gets expensive.
  • Quick back and forth. Use a smaller model for chatting.

Buy it if: you are writing something real, or running long jobs where a confident wrong answer would cost you.

Claude Sonnet 5

Anthropic · the daily default · $2 / $10

Good At

  • Almost everything, at a fair price. This is the one to default to.
  • Working on its own. The most capable Sonnet yet at planning a task and using a browser or terminal to finish it.
  • Daily coding, code review, and pulling information out of documents.

Not Good At

  • The absolute hardest problems. That is what Fable 5 and Opus are for.
  • Maximum effort settings. It emits more tokens per task, so pushed all the way it can actually cost more than Opus 4.8.

Buy it if: you want one model for most of your week and you do not want to think about it.

GPT-5.6 Sol

OpenAI · the flagship · $5 / $30

Good At

  • Speed on hard problems. It finishes complex tasks in roughly half the time Opus takes.
  • Agentic and terminal work. This is where it genuinely leads.
  • Coding that spans many files, tests, and follow-up fixes without losing the plot.

Not Good At

  • Being trusted without checking. It hallucinates noticeably more than Claude Opus 4.8, and at medium effort it has been caught skipping its own verification step, producing wrong answers that still pass the tests.
  • Cheap work. It is the priciest of the three GPT tiers.
  • Natural long-form writing. Claude reads better.

Buy it if: you need hard problems solved fast and you are going to review the output anyway.

GPT-5.6 Terra

OpenAI · the middle · $2.50 / $15

Good At

  • The broad middle of real work. Customer replies, reading documents, internal knowledge search, drafting at volume.
  • Short business copy. Hooks, subject lines, ad copy, calls to action.
  • Getting a lot done for the money. Roughly half the cost of the previous generation flagship.

Not Good At

  • Genuinely hard reasoning. That belongs with Sol.
  • Long-form voice work. It drifts off your tone over a long piece.

Buy it if: most of your work is business-as-usual and you want the best value per task.

GPT-5.6 Luna

OpenAI · the cheap one · $1 / $6

Good At

  • Simple, repetitive, high-volume tasks. Summaries, labeling, tagging, pulling fields out.
  • Rough first drafts you are going to rewrite anyway.
  • Cost. About a fifth of what Sol costs.

Not Good At

  • Long documents. This is the big one. Its ability to remember detail across a long file drops off sharply compared to Sol and Terra. Give it a big document and it will quietly miss things.
  • Anything with real stakes. Do not put it on work you will not check.
  • Nuanced writing. It is a workhorse, not a writer.

Buy it if: you have a pile of small boring tasks and you are paying too much to run them elsewhere.

Gemini 3.5 Flash

Google · the fast one · $1.50 / $9

Good At

  • Speed. It beats Google's own bigger model on coding and agent work while producing output several times faster.
  • Live search. The most reliable connection to current Google results of anything here, so it is the one for questions where being current matters.
  • Large documents and spreadsheets. It chews through big files without slowing down.
  • Google Workspace if your work already lives in Docs, Sheets, and Gmail.

Not Good At

  • Writing with personality. Gemini output tends to read flatter than Claude.
  • Being the top of the range. Gemini 3.5 Pro has been delayed repeatedly and is still not out, so Flash is your 3.5 option right now.

Buy it if: you want speed and current information, or your whole team already lives in Google.

Grok 4.3

xAI · the live one · $1.25 / $2.50

Good At

  • What is happening right now. It reads X live instead of a stale index, so on breaking events it is genuinely ahead.
  • Video. It takes video directly, up to five minutes, and can actually follow what happens in it.
  • Admitting it does not know. It scores unusually well at saying so instead of inventing an answer.
  • Cost. The cheapest output on this entire list.

Not Good At

  • Coding. It trails the others here, and xAI has not published the usual coding benchmarks for it.
  • Fast conversation. It takes about 12 seconds before it starts answering, which you feel immediately.
  • Very long inputs. Anything over 200,000 tokens of input is billed at double rate.

Buy it if: you need live information, you work with video, or you want frontier-adjacent quality for a fraction of the price.

Worth Knowing: The Free Lane

There is a second tier of models you can download and run for free, and they are close behind. The strongest right now are mostly Chinese, like DeepSeek, Qwen, Kimi, and GLM, plus Meta's Llama. They are called open weight, not open source, because the company hands you the finished model but not the training data or recipe. Not as strong as the eight above, but genuinely capable and free.

The MethodThe Three-Model Workflow

The workflow I keep seeing has three steps, and you do not use one model for all of them. You hand each part of the job to the model that is best at that part. Think first, then execute, then review.

Step 1

Think

Start with your smartest model. This is the one you use to plan the strategy, break the problem into pieces, and map out the approach. You want deep reasoning here, not speed. Give it room to think before anything gets built. In my flow this is a strong reasoning model like Claude's Fable 5.

Step 2

Execute

Now switch to your fastest model. The hard thinking is done, so this part is about doing the work quickly. Building, writing, coding, churning through the plan you just made. You do not need your deepest thinker for this, you need speed. A quick model like GPT-5.6 Sol does the heavy lifting fast.

Step 3

Review

Here is the step most people skip. Take the finished work and hand it to a different model to look for mistakes. Not the one that built it. A fresh model reads it cold and catches what the others missed. This single step has saved me more times than I can count.

Why It WorksWhy A Different Model To Review

Every model has different strengths. Every model also has different blind spots. That is the whole point of switching.

When one model builds something, it tends to be blind to its own mistakes. It made the same assumptions on the way in and on the way out. Ask it to check its own work and it will often tell you it looks great, because it is reading the problem the same way it did the first time.

A second, different model does not carry those same blind spots. It reads the work cold, with no attachment to how it got made, so it notices the gaps the first one talked itself past. It is the same reason a friend catches the typo you read past ten times. Fresh eyes see what tired eyes cannot.

The One Rule

The reviewer has to be a different model than the one that did the work. Same model, same blind spots. If you send the work back to the model that built it, you are asking it to catch a mistake it was already blind to. Switch the model and you switch the point of view.

In PracticeMy Lineup, By Job

Here is how the eight above map onto real jobs. This is a starting point, not a rulebook.

One Honest Note

Do not get too attached to the exact names. These change every few months, and something on this list will be replaced before the year is out. What stays true is the shape: smartest to think, fastest to build, a different one to review, cheapest for the boring stuff. Learn the shape and you can slot new models into it without redoing your whole setup.

The AnalogyThink Streaming Services

Here is how I explain it to people. AI models are like streaming services. Netflix, Hulu, all of them. If you only ever open one, you miss all the good stuff sitting on the others.

Nobody watches only Netflix and thinks they have seen everything worth watching. The good shows are spread across all of them. The skill is knowing which app to open for what you are in the mood for.

AI works the same way. The advantage is not owning the one perfect model. It is knowing which one to open for the job in front of you, and being able to switch in seconds.

The Takeaway

Stop hunting for the one best model. Start building a lineup. Your smartest model to think, your fastest to build, and a different one to review. That switch, from a favorite to a workflow, is what separates the people getting real work out of AI from the people still asking which one is best.

© 2026 Mariah Brunner. All rights reserved.