The 8 best AI agent frameworks

Best AI agent frameworks out of 60: LangGraph 54, Mastra and Microsoft Agent Framework 50. Tested on production, models, tracing, cost, hosting, proof.

Finlay Ekins
Co-founder · Updated · 14 min read
Share

What is an AI agent framework?

An AI agent framework is a code library for building agents that call models and tools, keep state, orchestrate multi-agent workflows and hand work to people when they need to.

LangGraph scores highest at 54/60, ahead of Mastra and Microsoft Agent Framework on 50. It is MIT licensed and names Klarna, Replit and Elastic as users (GitHub).

The six tests: production readiness, model freedom, observability, open licence and price, where it runs, and proof.

This page is for a developer or tech lead at an Australian business who is building an agent that runs inside the company's own systems. It ranks code frameworks only. For ready-made agents, not code frameworks, see the best AI agents.

Pick a coding agent instead? If the job is writing code faster rather than shipping an agent your business runs, a coding agent such as Claude Code, Codex, Cursor or Devin is the better fit. None of those is scored here.

The best AI agent frameworks, ranked

  1. LangGraph, 54/60: best for production control
  2. Mastra, 50/60: best for TypeScript teams
  3. Microsoft Agent Framework, 50/60: best for Azure shops
  4. Google ADK, 47/60: best for Google Cloud
  5. CrewAI, 41/60: best for multi-agent crews on a free plan
  6. Pydantic AI, 40/60: best for observability in Python
  7. OpenAI Agents SDK, 32/60: best for a lightweight start
  8. Claude Agent SDK, 23/60: best for teams that only run Claude

How each one scores

Points by test, scored 3 October 2026
Points for 8 frameworks on 6 tests, with the total out of 60.
LangGraphMastraMicrosoft Agent FrameworkGoogle ADKCrewAIPydantic AIOpenAI Agents SDKClaude Agent SDK
Production10 points
10 of 1010 of 1010 of 1010 of 1010 of 1010 of 106 of 106 of 10
Models10 points
7 of 107 of 1010 of 1010 of 106 of 107 of 107 of 103 of 10
Observability10 points
7 of 107 of 1010 of 107 of 104 of 1010 of 104 of 100 of 10
Cost10 points
10 of 1010 of 106 of 106 of 108 of 108 of 106 of 100 of 10
Deploy10 points
10 of 1010 of 1010 of 1010 of 1010 of 105 of 105 of 1010 of 10
Proof10 points
10 of 106 of 104 of 104 of 103 of 100 of 104 of 104 of 10
Total scoreout of 60
54 of 601st50 of 602nd50 of 603rd47 of 604th41 of 605th40 of 606th32 of 607th23 of 608th
Each test is scored out of 10, and every test counts the same. The best score on each line is in bold, and a tie is bold for everyone who shares it. The reason and source for every score are on the cards below.

Mastra and Microsoft Agent Framework tie on 50. The tie-break is proof, and Mastra leads it 6 to 4. Without cost and proof, Microsoft Agent Framework leads on 40.

Who wins each test:

  • Production: six-way tie on 10. Only OpenAI Agents SDK and Claude Agent SDK fall short, on 6.
  • Models: Microsoft Agent Framework and Google ADK, both 10.
  • Observability: Microsoft Agent Framework and Pydantic AI, both 10.
  • Cost: LangGraph and Mastra, both 10.
  • Deploy: six-way tie on 10. Only Pydantic AI and OpenAI Agents SDK fall short, on 5.
  • Proof: LangGraph alone, 10.

1. LangGraph: best for production control

1
LangGraph
LangChain
Best for
production control
Production10/10
Models7/10
Observability7/10
Cost10/10
Deploy10/10
Proof10/10
Total score54/60

LangGraph is LangChain's open-source orchestration library for stateful agents.

Review. LangGraph scores 54/60, the only framework here with a full 10 on proof, and it also takes full marks on production, cost and deploy. Its README says agents "persist through failures and can run for extended periods, automatically resuming from exactly where they left off" (GitHub). The licence is plain: "LangGraph is an MIT-licensed open-source library and is free" (LangChain). It names its users, "including Klarna, Replit, Elastic, and more" (GitHub), and ships a JS/TS library alongside Python. Hosting runs through LangSmith, which lists "$0 / seat per month" for Developer, "$39 / seat per month" for Plus and "Self-hosted and hybrid deployment options" on Enterprise (LangChain pricing).

Scores. Production 10 (durable 4, human approval 3, memory 3). Models 7 (any provider 4, MCP 3, A2A agent-to-agent protocol 0). Observability 7 (tracing 4, evals 3, OpenTelemetry 0). Cost 10. Deploy 10. Proof 10.

Loses on models: not stated: A2A. It drops 3 points to Microsoft Agent Framework and Google ADK.

Loses on observability: not stated: OpenTelemetry export.

Watch: the managed deployment beta is "public beta for Plus and Cloud Enterprise plans in the US" (LangChain pricing). No Australian region is stated.

Free 30-minute call

Need the agent built and run, not only the framework?

A framework gets you a stateful agent loop. It does not know your business, wire into your systems or keep running in production on its own. Book a time, tell us the job and the systems it touches, and you leave the call with what we would build and what it would cost.

For Australian companies of 10 to 100 people. What we build runs on your own servers and accounts, and you own all of the IP.

Book a 30-minute call

2. Mastra: best for TypeScript teams

2
Mastra
Mastra
Best for
TypeScript teams
Production10/10
Models7/10
Observability7/10
Cost10/10
Deploy10/10
Proof6/10
Total score50/60

Mastra is a TypeScript agent framework with its own cloud.

Review. Mastra scores 50/60 as a TypeScript AI agent framework and matches LangGraph on production, cost and deploy. Its "model routing connects to 90+ providers" and its "core framework is open source under the Apache 2.0 license" (Mastra). Code in its ee/ folder is source-available, not open source. Pricing is published: Starter is "$0 / month" and Teams is "$250 / month", and teams can "Self host your Mastra projects" (Mastra pricing). Customer stories name Salesforce ("How Salesforce Built Their Harness for 100k Developers on Mastra"), Range and Sanity (Mastra).

Scores. Production 10. Models 7 (any provider 4, MCP 3, A2A 0). Observability 7 (tracing 4, evals 3, OpenTelemetry 0). Cost 10. Deploy 10. Proof 6 (named customers 6, languages 0).

Loses on proof: it is TypeScript only, so it misses the 4 points for two or more languages.

Loses on models: not stated: A2A.

3. Microsoft Agent Framework: best for Azure shops

Best for
Azure shops
Production10/10
Models10/10
Observability10/10
Cost6/10
Deploy10/10
Proof4/10
Total score50/60

Microsoft Agent Framework is Microsoft's successor to AutoGen and Semantic Kernel. In Microsoft's words, it "combines AutoGen's simple agent abstractions with Semantic Kernel's enterprise features".

Review. Microsoft Agent Framework scores 50/60 and is the only framework with full marks on production, models and observability together. It "Supports Microsoft Foundry, Anthropic, Azure OpenAI, OpenAI, Ollama, and more" (Microsoft Learn). Its README lists "Built-in OpenTelemetry integration for distributed tracing, monitoring, and debugging", checkpointing, human-in-the-loop and A2A (GitHub). It ships in .NET and Python under MIT, with Go in public preview.

Scores. Production 10. Models 10. Observability 10. Cost 6 (licence 6, published price 0, free hosted tier 0). Deploy 10. Proof 4 (named customers 0, languages 4).

Loses on cost: not stated: a hosted price or free tier. Foundry usage is billed separately.

Loses on proof: no named customers stated.

4. Google ADK: best for Google Cloud

4
Best for
Google Cloud
Production10/10
Models10/10
Observability7/10
Cost6/10
Deploy10/10
Proof4/10
Total score47/60

Google ADK is Google's Agent Development Kit.

Review. Google ADK scores 47/60 and ties Microsoft Agent Framework for the top models score. "While optimized for Gemini, ADK is model-agnostic, deployment-agnostic" (GitHub), and it supports MCP tools and A2A. It is "Available in Python, TypeScript, Go, Java, and Kotlin", the widest language spread in this set (ADK docs). It is Apache 2.0, deploys to Vertex AI Agent Engine or Cloud Run, and has a bundled evaluation sample.

Scores. Production 10. Models 10. Observability 7 (tracing 4, evals 3, OpenTelemetry 0). Cost 6 (licence 6). Deploy 10. Proof 4 (languages 4).

Loses on observability: not stated: OpenTelemetry export.

Loses on proof: no named customers stated.

Loses on cost: no published hosted price. Agent Engine is billed on Google Cloud.

5. CrewAI: best for multi-agent crews on a free plan

5
CrewAI
CrewAI
Best for
multi-agent crews on a free plan
Production10/10
Models6/10
Observability4/10
Cost8/10
Deploy10/10
Proof3/10
Total score41/60

CrewAI is a Python framework for orchestrating teams of agents, with a commercial platform called AMP.

Review. CrewAI scores 41/60 and covers production in full, with "checkpointing, async execution, and MCP/A2A support" (GitHub). It is MIT licensed, the free plan includes "50 workflow executions/month" (CrewAI pricing), and AMP can be deployed "on-premise or in the cloud". Its home page says it is "Used by 65% of the Fortune 500" (CrewAI).

Scores. Production 10. Models 6 (any provider 0, MCP 3, A2A 3). Observability 4 (tracing 4). Cost 8 (licence 6, free tier 2). Deploy 10. Proof 3.

Loses on proof: the Fortune 500 claim scored 3 of 6 because customers appear only as logo images, and it is Python only.

Loses on observability: not stated: evals and OpenTelemetry.

Loses on models: not stated: support for many model providers.

crewai.com returned 403 to a direct fetch, so it was read through a reader copy.

6. Pydantic AI: best for observability in Python

6
Pydantic AI
Pydantic
Best for
observability in Python
Production10/10
Models7/10
Observability10/10
Cost8/10
Deploy5/10
Proof0/10
Total score40/60

Pydantic AI is the agent framework from the team behind Pydantic.

Review. Pydantic AI scores 40/60 and ties for the top observability score. Its docs describe the instrumentation as "plain OpenTelemetry", which works with any backend a team already runs, and Pydantic Evals tests agents (Pydantic AI). It supports "Virtually every model and provider", runs durable execution on Temporal, DBOS and Prefect, and can "run for days on the engine you already operate". It is MIT licensed, and Logfire "has a free tier that needs no credit card".

Scores. Production 10. Models 7 (any provider 4, MCP 3, A2A 0). Observability 10. Cost 8 (licence 6, free tier 2). Deploy 5 (self-host 5). Proof 0.

Loses on deploy: no managed agent hosting. Logfire is observability, not hosting.

Loses on proof: no named customers stated, and it is Python only.

7. OpenAI Agents SDK: best for a lightweight start

Best for
a lightweight start
Production6/10
Models7/10
Observability4/10
Cost6/10
Deploy5/10
Proof4/10
Total score32/60

OpenAI Agents SDK is OpenAI's open-source agent library.

Review. OpenAI Agents SDK scores 32/60. It is MIT licensed and "provider-agnostic, supporting the OpenAI Responses and Chat Completions APIs, as well as 100+ other LLMs" (GitHub). It has built-in tracing, sessions, human-in-the-loop and MCP tools, and a JavaScript/TypeScript version.

Scores. Production 6 (human approval 3, memory 3). Models 7 (any provider 4, MCP 3). Observability 4 (tracing 4). Cost 6 (licence 6). Deploy 5 (self-host 5). Proof 4 (languages 4).

Loses on production: not stated: durable execution.

Loses on observability: not stated: evals or OpenTelemetry.

Loses on deploy: not stated: managed hosting.

8. Claude Agent SDK: best for teams that only run Claude

Best for
teams that only run Claude
Production6/10
Models3/10
Observability0/10
Cost0/10
Deploy10/10
Proof4/10
Total score23/60

Claude Agent SDK is Anthropic's library for "the same tools, agent loop, and context management that power Claude Code, programmable in Python and TypeScript".

Review. Claude Agent SDK scores 23/60. It takes full marks on deploy, with sessions in "an Anthropic-managed cloud sandbox or a self-hosted sandbox on your own infrastructure" (Claude docs). It supports tool approval and MCP.

Scores. Production 6 (human approval 3, memory 3). Models 3 (MCP 3). Observability 0. Cost 0. Deploy 10. Proof 4 (languages 4).

Loses on cost: it is not open source. "Use of the Claude Agent SDK is governed by Anthropic's Commercial Terms of Service" (Claude docs).

Loses on models: only Claude models are stated.

Loses on observability: no tracing, evals or OpenTelemetry are stated.

How they compare

All eight pause for human approval, connect MCP tools and run on your own infrastructure, so the grid shows only the questions where they differ.

8 frameworks, side by side

Checked 3 October 2026
QuestionLangGraphMastraMicrosoft Agent FrameworkGoogle ADKCrewAIPydantic AIOpenAI Agents SDKClaude Agent SDK
Resumes after a failure or restart?Its own pages say runs are durable or resume after a failureLangGraph: Yes.Mastra: Yes.Microsoft Agent Framework: Yes.Google ADK: Yes.CrewAI: Yes.Pydantic AI: Yes.OpenAI Agents SDK: Not stated.Claude Agent SDK: Not stated.
Resumes after a failure or restart? Each framework's answer, from its own pages.
FrameworkAnswer
LangGraphYes. Says runs are durable or resume after a failure.
MastraYes. Says runs are durable or resume after a failure.
Microsoft Agent FrameworkYes. Says runs are durable or resume after a failure.
Google ADKYes. Says runs are durable or resume after a failure.
CrewAIYes. Says runs are durable or resume after a failure.
Pydantic AIYes. Says runs are durable or resume after a failure.
OpenAI Agents SDKNot stated. Its pages read do not say.
Claude Agent SDKNot stated. Its pages read do not say.
Works with many model providers?Its own pages say it supports models from other providersLangGraph: Yes.Mastra: Yes.Microsoft Agent Framework: Yes.Google ADK: Yes.CrewAI: Not stated.Pydantic AI: Yes.OpenAI Agents SDK: Yes.Claude Agent SDK: Not stated.
Works with many model providers? Each framework's answer, from its own pages.
FrameworkAnswer
LangGraphYes. Says it works with many model providers.
MastraYes. Says it works with many model providers.
Microsoft Agent FrameworkYes. Says it works with many model providers.
Google ADKYes. Says it works with many model providers.
CrewAINot stated. Its pages read do not say.
Pydantic AIYes. Says it works with many model providers.
OpenAI Agents SDKYes. Says it works with many model providers.
Claude Agent SDKNot stated. Its pages read do not say.
Built-in tracing?Its own pages name tracing or observability built inLangGraph: Yes.Mastra: Yes.Microsoft Agent Framework: Yes.Google ADK: Yes.CrewAI: Yes.Pydantic AI: Yes.OpenAI Agents SDK: Yes.Claude Agent SDK: Not stated.
Built-in tracing? Each framework's answer, from its own pages.
FrameworkAnswer
LangGraphYes. Names built-in tracing.
MastraYes. Names built-in tracing.
Microsoft Agent FrameworkYes. Names built-in tracing.
Google ADKYes. Names built-in tracing.
CrewAIYes. Names built-in tracing.
Pydantic AIYes. Names built-in tracing.
OpenAI Agents SDKYes. Names built-in tracing.
Claude Agent SDKNot stated. Its pages read do not say.
MIT or Apache 2.0 licence?Its licence file or pages say MIT or Apache 2.0LangGraph: Yes.Mastra: Yes.Microsoft Agent Framework: Yes.Google ADK: Yes.CrewAI: Yes.Pydantic AI: Yes.OpenAI Agents SDK: Yes.Claude Agent SDK: No.
MIT or Apache 2.0 licence? Each framework's answer, from its own pages.
FrameworkAnswer
LangGraphYes. MIT licence.
MastraYes. Apache 2.0 licence (core).
Microsoft Agent FrameworkYes. MIT licence.
Google ADKYes. Apache 2.0 licence.
CrewAIYes. MIT licence.
Pydantic AIYes. MIT licence.
OpenAI Agents SDKYes. MIT licence.
Claude Agent SDKNo. Governed by Anthropic's Commercial Terms of Service.
Yes answersOut of 4 questions4 yes for LangGraph4 yes for Mastra4 yes for Microsoft Agent Framework4 yes for Google ADK3 yes for CrewAI4 yes for Pydantic AI3 yes for OpenAI Agents SDK0 yes for Claude Agent SDK
Open a question to read each answer in full, with a link to the page it came from. A dash means the framework's own pages do not say, or the question is not a yes or no for it.

What an AI agent framework costs

List prices, USD
What an AI agent framework costs, by framework: Framework licence, Free hosted plan, Cheapest paid plan.
FrameworkFramework licenceFree hosted planCheapest paid plan
LangGraphFree, MITUS$0 a seat a monthLangSmith Developer, 5k tracesUS$39 a seat a monthLangSmith Plus
MastraFree, Apache 2.0 coreUS$0 a monthStarterUS$250 a monthTeams
Microsoft Agent FrameworkFree, MITNot publishedNot publishedmodel and Foundry usage billed separately
Google ADKFree, Apache 2.0Not publishedNot publishedAgent Engine billed on Google Cloud
CrewAIFree, MITFree50 workflow executions a monthNot publishedEnterprise, custom
Pydantic AIFree, MITFree Logfire tier, no cardNot published
OpenAI Agents SDKFree, MITNot publishedNot publishedmodel usage billed separately
Claude Agent SDKFree to install, Anthropic Commercial TermsNot publishedNot publishedClaude usage billed separately
List prices in US dollars as each vendor publishes them, before tax, read on 3 October 2026. No framework vendor publishes Australian dollar prices. Not published means no figure appears on the pages read. Model usage is billed separately by the model provider in every case.

Where a framework stops

The docs read for all eight describe the same core: the loop, the tools and the state. None of them describes learning the company's context, connecting the agent to the company's own systems, or maintaining it once it is live. Those three jobs sit with whoever builds and runs the agent.

The practical test is to pilot a framework on one real, repeated task before committing. Finlay Ekins puts the bar this way: "When a product can do something 80% as well as you can, whilst saving you time, then it becomes a must-use" (Medium). A framework that clears that bar on one task earns the next one.

If you would rather not own those three jobs, tell us the agent on a free 30-minute call.

What to check before you pick

  • Your language. Mastra is TypeScript only. CrewAI and Pydantic AI are Python only. Google ADK covers five languages.
  • Your cloud. Microsoft Agent Framework hosts on Foundry, Google ADK on Vertex AI Agent Engine, and LangGraph on LangSmith (US-only beta).
  • Your models. Claude Agent SDK states Claude only. Every other framework here states support for other providers, except CrewAI, whose pages read make no any-model claim.
  • Your licence. Seven of eight are MIT or Apache 2.0. Claude Agent SDK is under Anthropic's Commercial Terms.
  • Your data residency. No vendor states an Australian hosting region on the pages read. Self-hosting is available for all eight.

How we scored

How the scores work
  • Production, 10 points. Runs in production: resumes after failure or restart (4), pauses for human approval (3), keeps memory or sessions across runs (3).
  • Models, 10 points. No model lock-in: works with many model providers (4), MCP tools (3), A2A agent-to-agent protocol (3).
  • Observability, 10 points. You can see what it did: built-in tracing (4), evals (3), OpenTelemetry export (3).
  • Cost, 10 points. Open and priced: MIT or Apache 2.0 licence (6), published price for the hosted tier (2), free hosted tier (2).
  • Deploy, 10 points. Where it runs: a managed hosting option from the vendor (5), runs on your own infrastructure (5).
  • Proof, 10 points. Proof and reach: named customers on its own pages (6), two or more languages supported (4).

Total score out of 60. Scored 3 October 2026.

Each framework was scored on a 60-point rubric written for the buyer before any scoring, 6 tests of 10 points each. Scores come from a script that checks every claim against a saved copy of the vendor's page. Scores are floors: "not stated" or a dash means the pages read do not say it. A feature named only in a docs site's navigation is credited, as with tracing and resume for Google ADK, and the same standard applies to OpenAI's docs.

Tie-break: equal totals are ordered by proof.

What would lead without cost and proof. Drop those two tests and Microsoft Agent Framework leads on 40, ahead of Google ADK on 37 and LangGraph and Mastra on 34. LangGraph's lead comes from its published price, its free tier and its named customers.

Eligibility. A named company, public docs and public source code. All eight pass.

Not scored. Coding agents (see above) and no-code builders (Dify, n8n, Flowise, Langflow) are a different job. AutoGen and Semantic Kernel are merged into Microsoft Agent Framework. LlamaIndex, smolagents, AWS Strands and AG2 appear in comparison articles but were not read for this round, so they are candidates for the next refresh.

Set selection. The shortlist was drawn from the comparisons that rank for this query, including Firecrawl, Let's Data Science, Langfuse and Arize. These were used to pick the set, not as evidence.

References

  1. langchain.com/langgraph (cited for LangGraph)
  2. github.com/langchain-ai/langgraph/blob/main/README.md (cited for LangGraph)
  3. langchain.com/pricing (cited for LangGraph)
  4. mastra.ai (cited for Mastra)
  5. mastra.ai/pricing (cited for Mastra)
  6. learn.microsoft.com/en-us/agent-framework/overview/agent-framework-overview (cited for Microsoft Agent Framework)
  7. github.com/microsoft/agent-framework/blob/main/README.md (cited for Microsoft Agent Framework)
  8. google.github.io/adk-docs (cited for Google ADK)
  9. github.com/google/adk-python/blob/main/README.md (cited for Google ADK)
  10. crewai.com (cited for CrewAI)
  11. github.com/crewAIInc/crewAI/blob/main/README.md (cited for CrewAI)
  12. crewai.com/pricing (cited for CrewAI)
  13. ai.pydantic.dev (cited for Pydantic AI)
  14. openai.github.io/openai-agents-python (cited for OpenAI Agents SDK)
  15. github.com/openai/openai-agents-python/blob/main/README.md (cited for OpenAI Agents SDK)
  16. docs.claude.com/en/docs/agent-sdk/overview (cited for Claude Agent SDK)
  17. medium.com/@finlayekins/staying-on-top-of-ai-in-2026-f0ca06217e9d
  18. firecrawl.dev/blog/best-open-source-agent-frameworks
  19. letsdatascience.com/blog/ai-agent-frameworks-compared
  20. langfuse.com/blog/2025-03-19-ai-agent-comparison
  21. arize.com/guides/ai-agent-handbook/agent-frameworks

Common questions

What is the best AI agent framework in 2026?

LangGraph scores highest at 54/60. It takes full marks on production, cost, deploy and proof, and it is built for stateful agent orchestration. Mastra and Microsoft Agent Framework follow on 50.

Which AI agent framework is best for TypeScript?

Mastra, the only TypeScript-only framework here, scores 50/60. LangGraph, Google ADK, OpenAI Agents SDK and Claude Agent SDK also ship TypeScript versions.

Is LangGraph better than CrewAI?

On this rubric LangGraph scores 54 and CrewAI 41. The gap comes from models, observability, a published paid price and named customers. Both score 10 on production and deploy.

Which framework works with any model?

LangGraph, Mastra, Microsoft Agent Framework, Google ADK, Pydantic AI and OpenAI Agents SDK all state support for many providers. Claude Agent SDK states Claude only on the pages read.

Are AI agent frameworks free?

Seven of the eight are free under MIT or Apache 2.0. Claude Agent SDK is free to install but governed by Anthropic's Commercial Terms. Hosting and model usage cost extra in every case.

Where to start
$2,500flat, AI Audit
  • Every AI opportunity in your business, mapped in 7 to 14 days
  • Ranked roadmap with a spec and ROI figure for each build
  • Go ahead with the build within 30 days of your report and the full $2,500 comes off it
How the audit works

Turn this into real leverage.

We map where AI pays back in your business and build the agents that get you there.

Book a Discovery Call