This is an implementation of GraphRAG as described in
https://arxiv.org/pdf/2404.16130
From Local to Global: A Graph RAG Approach to Query-Focused Summarization
Official implementation by the authors of the paper is available at:
https://github.com/microsoft/graphrag/
While I generally prefer utilizing and refining existing implementations, as re-implementation often isn't optimal, I decided to take a different approach after encountering several challenges with the official version.
- Lacks integration with popular frameworks like LangChain, LlamaIndex, etc.
- Limited to OpenAI and AzureOpenAI models, with no support for other providers.
Using an established foundation like LangChain offers numerous benefits. It abstracts various providers, whether related to LLMs, embeddings, vector stores, etc., allowing for easy component swapping without altering core logic or adding complex support. More importantly, a solid foundation like this lets you focus on the problem's core logic rather than reinventing the wheel.
LangChain also supports advanced features like batching and streaming, provided your components align with the framework’s guidelines. For instance, using chains (LCEL) allows you to take full advantage of these capabilities.
The APIs are designed to be modular and extensible. You can replace any component with your own implementation as long as it implements the required interface.
Given the nature of the domain, this is important for conducting experiments by swapping out various components.
pip install langchain-graphragThere are 2 projects in the repo:
This is the core library that implements the GraphRAG paper. It is built on top of the langchain library.
Below is a snippet taken from the simple-app to show the style of API
and extensibility offered by the library.
Almost all the components (classes/functions) can be replaced by your own implementations. The library is designed to be modular and extensible.
# Reload the vector Store that stores
# the entity name & description embeddings
entities_vector_store = ChromaVectorStore(
collection_name="entity_name_description",
persist_directory=str(vector_store_dir),
embedding_function=make_embedding_instance(
embedding_type=embedding_type,
model=embedding_model,
cache_dir=cache_dir,
),
)
# Build the Context Selector using the default
# components; You can supply the various components
# and achieve as much extensibility as you want
# Below builds the one using default components.
context_selector = ContextSelector.build_default(
entities_vector_store=entities_vector_store,
entities_top_k=10,
community_level=cast(CommunityLevel, level),
)
# Context Builder is responsible for taking the
# result of Context Selector & building the
# actual context to be inserted into the prompt
# Keeping these two separate further increases
# extensibility & maintainability
context_builder = ContextBuilder.build_default(
token_counter=TiktokenCounter(),
)
# load the artifacts
artifacts = load_artifacts(artifacts_dir)
# Make a langchain retriever that relies on
# context selection & builder
retriever = LocalSearchRetriever(
context_selector=context_selector,
context_builder=context_builder,
artifacts=artifacts,
)
# Build the LocalSearch object
local_search = LocalSearch(
prompt_builder=LocalSearchPromptBuilder(),
llm=make_llm_instance(llm_type, llm_model, cache_dir),
retriever=retriever,
)
# it's a callable that returns the chain
search_chain = local_search()
# you could invoke
# print(search_chain.invoke(query))
# or, you could stream
for chunk in search_chain.stream(query):
print(chunk, end="", flush=True)git clone https://github.com/ksachdeva/langchain-graphrag.gitDevcontainer will install all the dependencies
- Clone the repository
git clone https://github.com/ksachdeva/langchain-graphrag.git
cd langchain-graphrag- Install dependencies (requires Python 3.10+ and uv)
You can install uv using the standalone installers or from PyPI:
# On macOS and Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
# On Windows
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"# With pip
pip install uv
# Or pipx
pipx install uvIf installed via the standalone installer, you can update uv to the latest version:
uv self updateuv syncThis is a simple typer based CLI app.
In terms of configuration it is limited by the number of command line options exposed.
That said, the way core library is written you can easily replace any component by your own implementation i.e. your choice of LLM, embedding models etc. Even some of the classes as long as they implement the required interface.
Note:
Copy examples/simple-app/.env.example to examples/simple-app/.env and fill in your API keys and provider settings.
cp examples/simple-app/.env.example examples/simple-app/.envEdit examples/simple-app/.env to choose your provider and set API keys:
# Provider — pick one: openai | azure_openai | ollama
LLM_TYPE=openai
LLM_MODEL=gpt-4o
EMBEDDING_TYPE=openai
EMBEDDING_MODEL=text-embedding-3-small# Index using the provider configured in .env
uv run poe index
# To see all indexer options
uv run poe indexer-helpuv run poe global-search --query "What are the top themes in this story?"
# To see all options
uv run poe query-helpuv run poe local-search --query "Who is Scrooge, and what are his main relationships?"See examples/simple-app/README.md for more details.
The project includes several convenient poe tasks (see poe.toml for complete list):
# App commands
uv run poe app-help # General help
uv run poe indexer-help # Indexer help
uv run poe query-help # Query help
# Run (provider configured via examples/simple-app/.env)
uv run poe index # Index input data
uv run poe report # Generate reports (requires prior indexing)
uv run poe global-search --query "your question"
uv run poe local-search --query "your question"
# Development
uv run poe test # Run tests
uv run poe lint # Check code quality
uv run poe format # Format code
uv run poe typecheck # Type checking
uv run poe docs-serve # Serve documentation locally# 1. Setup
uv sync
# 2. Configure provider and API keys
cp examples/simple-app/.env.example examples/simple-app/.env
# Edit examples/simple-app/.env — fill in API keys, set LLM_TYPE
# 3. Index and search
uv run poe index
uv run poe global-search --query "What are the themes?"
uv run poe local-search --query "Who is the main character?"
# 4. Development (optional)
uv run poe test && uv run poe lint # Test and check codeThe state of the library is far from complete.
Here are some of the things that need to be done to make it more useful:
- Add more guides
- Document the APIs
- Add more tests