Why LLM frameworks exist
A language model such as GPT or Claude is trained on a huge amount of public internet text, so it can answer general questions well. But real applications need more than a single model call:
- The model does not know your private data. Ask a chatbot for a fictional clothing store’s refund policy on t-shirts and it cannot answer, because the policy document was never part of its training.
- Real tasks take several steps: load data, clean it, call a model, call another service, format the result.
- Useful assistants need to connect to tools such as databases, search engines, and internal APIs.
You could write all of this by hand in plain Python or JavaScript, and for small projects that is a perfectly good choice. Frameworks save time by giving you ready-made building blocks for the repetitive parts: loading files, splitting text, talking to vector databases, calling tools, and managing conversation state.
The first fix for the “model does not know my data” problem is called Retrieval-Augmented Generation (RAG): the system retrieves relevant passages from your documents and gives them to the model along with the question. If you are new to the idea, our guide to RAG explained covers it step by step.
LlamaIndex: the framework for your data
LlamaIndex is an open-source framework for connecting your own data to large language models. Its center of gravity is knowledge: getting your documents into a form a model can search, and getting accurate answers back out.
The three steps LlamaIndex is built around
- Ingest. Connect to your data sources: PDFs, Word files, databases, APIs, Notion, and many more. Hundreds of connectors and integrations exist for this.
- Index. Turn the data into searchable structures, typically vector embeddings plus metadata, and store them in a vector database of your choice.
- Query. Ask a question through a query interface that finds the relevant content and returns a knowledge-grounded answer.
LlamaIndex works with unstructured data (documents, text), structured data (SQL tables), and semi-structured data (JSON, APIs). It also provides more advanced retrieval tools, such as query routers that decide which data source is best placed to answer a question.
A minimal example
Install it with pip install llama-index, put some files in a data folder, and set your OPENAI_API_KEY environment variable (the default setup uses OpenAI for embeddings and answers):
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
# Ingest: load every file in the "data" folder
documents = SimpleDirectoryReader("data").load_data()
# Index: build a searchable vector index
index = VectorStoreIndex.from_documents(documents)
# Query: ask a question over your documents
query_engine = index.as_query_engine()
response = query_engine.query("What is the refund policy for t-shirts?")
print(response)
Those few lines cover ingestion, indexing, retrieval, and answer generation.
When LlamaIndex is a great fit
- An internal knowledge base for a company
- A question-and-answer bot over legal, medical, or technical documents
- Any project where accurate answers from a specific body of information are the main goal
Note: LlamaIndex is no longer only about RAG. It now also offers agents and event-driven workflows, and a separate cloud platform (LlamaParse) for parsing complicated documents. Its identity is still “the data framework”, but the older idea that it can only do search is out of date.
LangChain: the framework for building LLM apps
LangChain is a general-purpose, open-source framework for building applications powered by language models. Where LlamaIndex is about knowledge, LangChain is about building things that do something. It gives you a modular toolbox, with a large ecosystem of integrations for model providers, vector stores, tools, and services, and you snap pieces together.
Two shapes of application are common:
- Chains: a fixed sequence of steps where a model is one of the components. Load data, preprocess, ask the model to extract or summarize something, pass the result on.
- Agents: systems where the model itself decides which tool to use next and when the job is done.
LangChain also provides ready-made classes for typical RAG steps (text loaders, splitters, vector store connectors), so it can build RAG systems too, just with a broader scope than LlamaIndex.
A minimal agent example
Install with pip install -U langchain "langchain[openai]". Current LangChain versions use create_agent as the standard way to build an agent:
from langchain.agents import create_agent
def get_order_status(order_id: str) -> str:
"""Look up the status of an order by its ID."""
# In a real app, query your database or API here
return f"Order {order_id} is out for delivery."
agent = create_agent(
model="openai:gpt-4o", # use any chat model you have access to
tools=[get_order_status],
system_prompt="You are a helpful customer support assistant.",
)
result = agent.invoke(
{"messages": [{"role": "user", "content": "Where is order 1042?"}]}
)
print(result["messages"][-1].content)
The model reads the question, decides it needs the get_order_status tool, calls it, and writes the final reply from the result. Framework APIs change quickly, so check the official documentation if a name or import differs in your installed version.
When LangChain is a great fit
- A support agent that can look up an order, process a refund, and send a confirmation email
- An assistant that combines web search with a code interpreter to solve a problem
- Any app that needs flexibility, many integrations, and custom multi-step logic
Workflows vs agents
To understand where LangGraph fits, it helps to separate two kinds of LLM-powered applications, a distinction popularised by Anthropic’s guide to building effective agents.
Workflows are ordinary software where the steps follow a predefined path, and a model is one of the components. You load data, preprocess it, call the model to extract something, then continue. The model does not decide what happens next. Your code does.
Agents are systems where the model has some autonomy. It plans, chooses tools, checks results, and retries when something fails, much like a developer who runs code, sees an error, fixes it, and runs again.
A worked example
Suppose a customer of an online clothing store types: “I want to exchange my t-shirt for a different item.” That cannot be answered in one shot. The system has to:
- Check the return policy for that item (searching policy documents)
- Ask for the order ID, then check the order details in the database to confirm eligibility
- Ask what the customer wants instead, then check availability of the new item (an inventory API)
- Place the new order and print a shipping label for the old item
Some steps need clarification from the user, and some may fail. If the requested colour is out of stock, the system should try another option. Knowledge (documents, databases) and tools (APIs) are combined under the model’s guidance. That is an agentic workflow, and it needs more than a simple linear chain.
LangGraph: stateful, multi-step agents
LangGraph is a framework for orchestrating complex agent workflows as a graph. Each step is a node. Edges between nodes can be conditional, so after a model call the flow may go to step A, step B, another model, or loop back to where it started. The graph also carries state (the conversation so far, intermediate results, and so on), which makes memory, loops, and retries natural.
In practice, a LangGraph project might define a chatbot node and a tool node. The model calls a tool such as a stock-price lookup, gets a result, and can call it again as many times as needed before giving a final answer. For instance, “buy 20 shares” can trigger a price lookup and then a total calculation.
The official LangChain documentation positions the layers like this: use LangChain’s create_agent when you want a customizable agent quickly, and use LangGraph, the lower-level orchestration framework, when you need finer control, such as combining deterministic steps with agentic ones, human approval steps, or custom loops. LangChain agents are in fact built on the same foundation, so the two work together rather than compete.
LangSmith: debugging and monitoring
LangSmith is different from the others. It is not a framework for building the app. It is a platform for observing it: debugging, testing, and monitoring once it is running.
With tracing turned on, you can see every individual call in a request: the vector database lookup, the model call, the exact input and output of each step, how many tokens were used, and how long it took. That matters most in production, where you need to watch quality, latency, and token costs, and investigate failures. Tracing is switched on by setting environment variables (for example LANGSMITH_TRACING=true and your API key). Despite the name, it can be used with applications not built on LangChain.
Side-by-side comparison
LangChain vs LlamaIndex
| Feature | LlamaIndex | LangChain |
|---|---|---|
| Primary focus | Connecting your data to LLMs: ingestion, indexing, retrieval | General-purpose building of LLM apps, chains, and agents |
| Data handling | Deep tooling for parsing, structuring, and querying private data | Loaders, splitters, and vector store integrations for common cases |
| Best for | RAG, document Q&A, knowledge bases | Agents, tool use, multi-step logic |
| Customization | Strong for data pipelines and retrieval strategies | Highly modular, easy to combine tools and components |
| Learning curve | Often considered quicker to start for a data use case | Broader, so more concepts to learn |
| Integrations | Hundreds of connectors and integration packages | Very large ecosystem of model, tool, and store integrations |
| Languages | Python and TypeScript | Python and JavaScript/TypeScript |
LangChain vs LangGraph vs LangSmith
| LangChain | LangGraph | LangSmith | |
|---|---|---|---|
| What it is | Toolkit for LLM apps, chains, and agents | Low-level framework for stateful, graph-based agent workflows | Platform for tracing, debugging, evaluation, and monitoring |
| Use it to | Build apps and standard agents quickly | Control complex flows with loops, retries, branches, and memory | See what your app is doing and fix or measure it |
| Open source | Yes | Yes | Hosted platform (with a free tier) |
| Typical use case | Chatbots, simple RAG, tool-using agents | Multi-step agents, approval flows, complex automation | Production monitoring, token cost tracking, evaluation |
A note on speed claims: you will often read that LlamaIndex retrieves faster. Its indexing and retrieval tooling is its main focus, so it offers more built-in options here, but real-world speed depends mostly on your vector database, chunking, model, and infrastructure. Benchmark on your own data before deciding.
Using LlamaIndex and LangChain together
You do not have to pick only one. A common production pattern is to give each tool the job it is best at:
- LlamaIndex is the data brain. It ingests and indexes your documents and answers knowledge questions.
- LangChain (or LangGraph) is the action brain. It handles reasoning, planning, and tool use, and treats the LlamaIndex query engine as just another tool.
User question
|
v
LangChain / LangGraph agent (reasoning, planning, tool use)
| | |
v v v
LlamaIndex Order API Refund API
(documents) (database) (payments)
|
v
Answer written by the LLM
Here is a small example of that idea. The LlamaIndex query engine is wrapped as a LangChain tool, next to an action tool:
from langchain.agents import create_agent
from langchain.tools import tool
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
# Data brain: index the policy documents with LlamaIndex
documents = SimpleDirectoryReader("data").load_data()
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine()
@tool
def search_policies(question: str) -> str:
"""Answer questions about company policies using the policy documents."""
return str(query_engine.query(question))
@tool
def get_order_status(order_id: str) -> str:
"""Look up the status of an order by its ID."""
return f"Order {order_id} is out for delivery." # replace with a real lookup
# Action brain: a LangChain agent that can use both tools
agent = create_agent(
model="openai:gpt-4o",
tools=[search_policies, get_order_status],
system_prompt="You are a support assistant. Use the tools to answer.",
)
result = agent.invoke(
{"messages": [{"role": "user", "content": "Can I return order 1042 for a refund?"}]}
)
print(result["messages"][-1].content)
With this setup the agent can look up the order, check the refund policy in your documents, and combine both into one answer. Test it with your own installed versions, as package APIs evolve.
What this means for mobile app developers
If you build with Flutter, React Native, Swift, or Kotlin, one point matters more than any framework comparison: these frameworks run on a server, not inside your app.
Mobile app --HTTPS--> Your backend (Python or Node) --> LlamaIndex / LangChain / LangGraph
--> Vector database, LLM, other APIs
--> LangSmith (tracing)
Reasons for this design:
- Security. API keys placed inside an app can be extracted. Keep them on the server.
- Performance and battery. Embeddings, vector search, and multi-step agent loops belong on a server.
- Easy updates. Change prompts, tools, models, or frameworks on the backend without an app store release.
- Cost control. Add authentication, rate limits, and usage caps in one place.
Both ecosystems have JavaScript/TypeScript versions, so a React Native team can keep a single language across the app and a Node backend. Python remains the most common choice for the AI layer.
Practical tips for AI features in apps
- Stream responses. Agents can take several seconds. Show text as it arrives, or a clear progress state, instead of a frozen screen.
- Design for failure. Tools time out and networks drop. Handle errors with clear messages and retry options.
- Show sources for knowledge answers so users can trust them.
- Require confirmation for actions. If an agent can cancel an order or send money, ask the user to approve first.
- Watch your token costs. Agents can make many model calls per request. Trace them and set limits.
Which one should you choose?
| Your app needs to… | Start with |
|---|---|
| Answer questions from your help docs, PDFs, or knowledge base | LlamaIndex |
| Call tools and APIs, such as checking orders or sending emails | LangChain |
| Run a complex multi-step flow with loops, retries, memory, or human approval | LangGraph |
| Monitor quality, latency, and token costs in production | LangSmith |
| Both know your data and take actions | LlamaIndex for the data plus LangChain or LangGraph for the agent |
A simple rule of thumb: if your project is mainly about RAG, lean toward LlamaIndex. If you are building agents that act, lean toward LangChain and, for complex control, LangGraph. For the most capable applications, combine them.
Do you need a framework at all? For a small feature, such as one chatbot over a few PDFs, plain Python with a vector database and one model API is often enough, and easier to understand and maintain. Frameworks pay off as complexity grows.
Common mistakes to avoid
- Picking a framework before defining the feature. Decide whether your feature needs to know things, do things, or both.
- Using an agent when a workflow would do. If the steps are always the same, a fixed chain is cheaper, faster, and more predictable than a free-roaming agent.
- Putting AI keys or logic in the mobile app. Always go through your own backend.
- Skipping observability. Without tracing, you cannot tell why an agent gave a wrong answer or why your bill jumped.
- Copying old tutorials blindly. These libraries change quickly. Check the current documentation for your installed version.
Frequently asked questions
Is LlamaIndex better than LangChain?
Neither is better in general. LlamaIndex is strongest for connecting private data to models and for RAG. LangChain is strongest for general application building and agents. They can also be used together.
Can LangChain do RAG?
Yes. It includes document loaders, text splitters, and vector store integrations. LlamaIndex simply goes deeper on the data and retrieval side.
Can LlamaIndex build agents?
Yes. It now includes agent and workflow features, and it can use a RAG pipeline as one tool among many.
What is the difference between LangChain and LangGraph?
LangChain is a toolkit for building LLM apps and standard agents. LangGraph is a lower-level framework for stateful, graph-based workflows with loops, branches, and retries. LangChain agents are built on top of the same foundation.
What does LangSmith do?
It traces, debugs, evaluates, and monitors LLM applications, showing each call, its inputs and outputs, token usage, and latency.
Do I have to use LangChain or LlamaIndex to build an AI app?
No. You can call a model API directly and write the rest yourself. Frameworks help most when your app grows in complexity.
Can I use these frameworks directly in a mobile app?
It is better not to. Run them on a backend server and let your app call it over HTTPS. That keeps your keys safe and your app fast.
Which framework is best for beginners?
For a first project over your own documents, LlamaIndex is often the quickest to get working. For a first tool-using assistant, LangChain’s create_agent is a good starting point.
Conclusion
The four names are easier to understand once you see the job each one does. LlamaIndex helps your app know your data. LangChain helps it act. LangGraph gives you precise control when the actions become a complicated, stateful process. LangSmith lets you see what is going on so you can improve it.
Start from what your feature needs, keep the AI logic on a backend, and add pieces only when the project calls for them. Build a small prototype with one framework first, try it on real questions from your users, and expand from there.