LLMs

Sep 2026 Link your models

Best for

  • Writing
    Anthropic logoClaude Fable 5.1
    Anthropic logoClaude Sonnet 5Anthropic logoClaude Opus 5.5
  • Conversation
    Google logoGemini 3.8 Flash
    Anthropic logoClaude Sonnet 5Anthropic logoClaude Opus 5.5
  • Coding
    Moonshot AI logoKimi K3
    Alibaba logoQwen3.8 MaxAnthropic logoClaude Opus 5.5
  • Reasoning & math
    OpenAI logoGPT-6 Astra
    Moonshot AI logoKimi K3Anthropic logoClaude Fable 5.1
  • Web research
    Anthropic logoClaude Opus 5.5
    Anthropic logoClaude Fable 5.1OpenAI logoGPT-6 Astra
  • Data & analysis
    Anthropic logoClaude Fable 5.1
    OpenAI logoGPT-6 AstraAnthropic logoClaude Opus 5.5
  • Long documents
    OpenAI logoGPT-6 Astra
    Meta logoMuse Spark 1.3Moonshot AI logoKimi K3
  • Images & PDFs
    Google logoGemini 3.8 Flash
    Google logoGemini 3.1 Pro (Preview)Anthropic logoClaude Fable 5.1
  • Languages
    Google logoGemini 3.8 Flash
    Google logoGemini 3.1 Pro (Preview)Alibaba logoQwen3.8 Max
  • Quick replies
    OpenAI logoGPT-5.6 Luna
    Google logoGemini 3.8 FlashDeepSeek logoDeepSeek V4 Pro

Models

  • GPT-6 Astra

    OpenAI · Sep 2026
    • Reasoning & math
    • Long documents
    • Data & analysis
    $10 / $50 per 1M · 1.1M context
    Input: Text · Image
    • Computer and browser use
    • Long-horizon agents
    • Scientific reasoning
    • Hard math
    • Agentic coding
    • Deep web research
    AA Index
    53
    Arena Elo
    1480
    GPQA
    96%
    HLE+tools
    57.2%
    • Raise reasoning effort for hard problems
    • Use fast mode (2x price) for latency
    • Cache long prompts: $1 per 1M cached
    FinCraftly routes coding & reasoning & math hereLink
  • GPT-5.6 Luna

    OpenAI · Jul 2026
    • Quick replies
    $0.20 / $1.20 per 1M · 1.1M context
    Input: Text · Image
    • High-volume automation
    • Classification and extraction
    • Chat support bots
    • Cheap sub-agents
    • Summarization
    AA Index
    37
    GPQA
    92.3%
    SWE Pro
    62.7%
    TB 2.1
    84.7%
    • Use for bulk, cost-sensitive jobs
    • Set reasoning effort none for speed
    • Escalate hard cases to Astra
    FinCraftly routes quick replies & conversation hereLink
  • Claude Opus 5.5

    Anthropic · Sep 2026
    • Web research
    • Writing
    • Conversation
    $4 / $20 per 1M · 1M context
    Input: Text · Image
    • Agentic coding
    • Long-running agents
    • Professional knowledge work
    • Long-form writing
    • Computer use
    • Document analysis
    HLE+tools
    67.7%
    TB 4.0
    66.4%
    • Default pick for most agent workloads
    • Use xhigh effort for hardest tasks
    • Fast mode: 2x price, faster output
    FinCraftly routes writing & conversation hereLink
  • Claude Fable 5.1

    Anthropic · Sep 2026
    • Writing
    • Data & analysis
    • Web research
    $10 / $50 per 1M · 1M context
    Input: Text · Image
    • Multi-hour agent runs
    • Scientific research
    • Hard reasoning
    • Agentic coding
    • Complex analysis
    AA Index
    53
    Arena Elo
    1498
    HLE+tools
    65.6%
    TB 4.0
    55.8%
    • Reserve for the hardest, longest tasks
    • Cache reads $0.25/1M cut agent costs
    • Try Opus 5.5 first; similar on most work
    FinCraftly routes writing & coding hereLink
  • Claude Sonnet 5

    Anthropic · Jun 2026
    • Writing
    • Conversation
    $2 / $10 per 1M · 1M context
    Input: Text · Image
    • Everyday coding
    • Business writing
    • Customer-facing chat
    • Computer use
    • Agent workflows
    AA Index
    38
    HLE+tools
    57.4%
    HLE
    43.2%
    SWE Pro
    63.2%
    • Best value Claude for daily work
    • Lower effort for faster replies
    • Upgrade to Opus 5.5 for long agents
    FinCraftly routes writing & conversation hereLink
  • Gemini 3.8 Flash

    Google · Sep 2026
    • Conversation
    • Images & PDFs
    • Languages
    $0.75 / $3.75 per 1M · 1M context
    Input: Text · Image · Audio · Video · Promotional until 2026-12-31; $1.50 / $7.50 from 2027-01-01
    • Fast agentic coding
    • Multimodal understanding
    • Video and audio input
    • High-volume apps
    • Chart reasoning
    AA Index
    41
    Arena Elo
    1493
    HLE
    45.4%
    SWE Pro
    61.6%
    • Very fast; good default for agents
    • Use high thinking for hard tasks
    • Price doubles on 2027-01-01
    FinCraftly routes conversation & long documents hereLink
  • Gemini 3.1 Pro (Preview)

    Google · Feb 2026
    • Images & PDFs
    • Languages
    $2 / $12 per 1M · 1M context
    Input: Text · Image · Audio · Video · Prompts over 200K tokens: $4 / $18
    • Multimodal analysis
    • Long documents
    • Visual reasoning
    • Creative coding
    • Research synthesis
    AA Index
    30
    Arena Elo
    1487
    GPQA
    94.3%
    HLE+tools
    51.4%
    • Keep prompts under 200K for base price
    • Strong at video and image inputs
    • Try 3.8 Flash first for speed
    FinCraftly routes long documents & images & pdfs hereLink
  • Grok 4.7

    xAI · Sep 2026
    • Conversation
    • Coding
    $2 / $6 per 1M · 500K context
    Input: Text · Image · Prompts over 200K tokens: $4 / $12
    • Coding
    • Engineering problems
    • Professional knowledge work
    • Real-time X context
    • Cost-efficient reasoning
    AA Index
    46
    TB 4.0
    38%
    • Pricing doubles above 200K prompt tokens
    • Use xhigh effort on hard tasks
    • Fast variant: 2x price, 2x speed
    FinCraftly routes conversation & coding hereLink
  • Muse Spark 1.3

    Meta · Sep 2026
    • Long documents
    $1.25 / $4.25 per 1M · 1M context
    Input: Text · Image · Audio · Video · Pdf
    • Multi-agent workflows
    • Agentic coding
    • Long-context recall
    • Multimodal input
    • Low-cost frontier work
    AA Index
    48
    Arena Elo
    1493
    GPQA
    94%
    TB 2.1
    88.8%
    • Very low cost per task
    • Cached input only $0.15/1M
    • Access via Meta Model API
  • DeepSeek V4 Pro

    DeepSeek · Apr 2026
    • Quick replies
    $1.32 / $3.96 per 1M · 1M context · Open weights
    Input: Text · Peak rate; off-peak (most hours) is $0.66 / $1.98
    • Math and reasoning
    • Competitive coding
    • Self-hosting
    • Budget API use
    • Long documents
    AA Index
    36
    GPQA
    90.1%
    HLE+tools
    48.2%
    SWE Pro
    55.4%
    • Off-peak hours cost half
    • Weights are MIT-licensed
    • No image input; use Flash for vision
    FinCraftly routes quick replies & coding hereLink
  • Qwen3.8 Max

    Alibaba · Sep 2026
    • Coding
    • Languages
    $2 / $6 per 1M · 1M context
    Input: Text · Image · Video
    • Multi-step coding projects
    • Tool orchestration
    • Document parsing
    • Chart reasoning
    • Multilingual tasks
    AA Index
    45
    Arena Elo
    1481
    GPQA
    92.6%
    SWE Pro
    67.7%
    • Good for long autonomous coding runs
    • Use implicit caching to cut cost
    • Strong for Chinese and multilingual
    FinCraftly routes coding & long documents hereLink
  • Kimi K3

    Moonshot AI · Jul 2026
    • Coding
    • Reasoning & math
    • Long documents
    $3 / $15 per 1M · 1M context · Open weights
    Input: Text · Image · Video
    • Agentic web research
    • Large-repo coding
    • Frontend and UI code
    • Knowledge work
    • Self-hosting
    AA Index
    44
    Arena Elo
    1485
    GPQA
    93.5%
    HLE+tools
    56%
    • Top open model for browsing agents
    • Cache hits cost $0.30/1M
    • Self-host with vLLM or SGLang
    FinCraftly routes coding & reasoning & math hereLink
  • GLM-5.3

    Z.ai · Aug 2026
    • Conversation
    • Coding
    $1.40 / $4.40 per 1M · 1M context · Open weights
    Input: Text
    • Agentic coding
    • Automation workflows
    • Security analysis
    • Self-hosting
    • Budget reasoning
    AA Index
    45
    Arena Elo
    1483
    • Reasoning is always on
    • Flash variant for cheap bulk work
    • Check custom licence before hosting
  • Mistral Medium 3.5

    Mistral AI · May 2026
    • Languages
    • Writing
    $1.50 / $7.50 per 1M · 262K context · Open weights
    Input: Text · Image
    • European data residency
    • Self-hosted agents
    • Coding with Vibe
    • Structured outputs
    • Multilingual business
    AA Index
    14
    SWE Verified
    77.6%
    • Runs on as few as four GPUs
    • Set reasoning effort per request
    • Good fit for EU compliance needs
    FinCraftly routes languages & writing hereLink
  • Command A+

    Cohere · May 2026
    • Languages
    • Quick replies
    per 1M · 128K context · Open weights
    Input: Text · Image
    • Enterprise RAG
    • Private deployment
    • Multilingual (48 languages)
    • Document processing
    • Tool use
    AA Index
    13
    • Runs on 2x H100 or 1x B200
    • Apache 2.0: fully self-hostable
    • Pair with Cohere Rerank for RAG

Benchmarks

ModelGPQAHLEHLE+toolsSWE ProSWE VerifiedTB 4.0TB 2.1ARC-AGI-2Arena EloAA Index
Anthropic logoClaude Fable 5.165.6%55.8%90%149853
OpenAI logoGPT-6 Astra96%57.2%57.9%95%148053
Meta logoMuse Spark 1.394%88.8%149348
xAI logoGrok 4.738%46
Z.ai logoGLM-5.3148345
Alibaba logoQwen3.8 Max92.6%67.7%86.6%148145
Moonshot AI logoKimi K393.5%43.5%56%88.3%148544
Google logoGemini 3.8 Flash45.4%61.6%90.8%149341
Anthropic logoClaude Sonnet 543.2%57.4%63.2%80.4%38
OpenAI logoGPT-5.6 Luna92.3%62.7%84.7%37
DeepSeek logoDeepSeek V4 Pro90.1%48.2%55.4%80.6%36
Google logoGemini 3.1 Pro (Preview)94.3%44.4%51.4%54.2%80.6%77.1%148730
Mistral AI logoMistral Medium 3.577.6%14
Cohere logoCommand A+13
Anthropic logoClaude Opus 5.567.7%66.4%
Sources