If you want a private, zero-cost way to ask questions of your own PDFs (contracts, manuals, internal reports), a fully local Retrieval‑Augmented Generation (RAG) stack is an excellent choice. This guide walks you through a tested 2026 workflow that uses Ollama to run local models, ChromaDB as a persistent vector store, and LangChain to orchestrate ingestion, retrieval, and generation. Everything runs on your machine — no API keys, no cloud costs, and nothing leaves your network. The instructions below are practical and runnable end-to-end.
ollama serve or the installed systemd unit). (ollama.readthedocs.io) ollama pull flow downloads models into your local Ollama store so inference is free and offline. (ollama.com)Create a virtualenv and install the necessary packages: - Python packages (example): - pip install --upgrade pip - pip install langchain langchain-ollama langchain-chroma langchain-community chromadb pypdf pymupdf
Notes: package names and exact versions change quickly; LangChain provides maintained integration packages for Ollama and Chroma. Use the latest compatible versions. (reference.langchain.com)
High-level flow: load PDFs → split into chunks → embed chunks via Ollama → store vectors in Chroma (persistent).
Example (ingest.py — annotated):
from pathlib import Path
import hashlib
from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_ollama import OllamaEmbeddings
from langchain_chroma import Chroma
import chromadb
# Config
OLLAMA_BASE = "http://localhost:11434"
EMBED_MODEL = "nomic-embed-text"
CHROMA_DIR = "./chroma_data"
COLLECTION = "pdf_docs"
# Loader + splitter
loader = PyPDFLoader # uses pypdf/pymupdf under the hood
splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
# Embeddings via local Ollama
embeddings = OllamaEmbeddings(model=EMBED_MODEL, base_url=OLLAMA_BASE)
# Persistent Chroma client + LangChain wrapper
client = chromadb.PersistentClient(path=CHROMA_DIR)
vectordb = Chroma(client=client, collection_name=COLLECTION, embedding_function=embeddings)
def stable_id(path: str, idx: int) -> str:
return hashlib.sha256(f"{path}::chunk_{idx}".encode()).hexdigest()[:16]
def ingest_pdf(path: str):
docs = loader(path).load()
chunks = splitter.split_documents(docs)
for i, c in enumerate(chunks):
c.metadata.update({"source": Path(path).name, "chunk_index": i, "doc_id": stable_id(path, i)})
vectordb.add_documents(chunks, ids=[c.metadata["doc_id"] for c in chunks])
This pattern — PyPDFLoader + RecursiveCharacterTextSplitter + OllamaEmbeddings + Chroma — is the standard local-RAG ingestion pipeline used across recent tutorials and examples. Persist the Chroma data directory so re-runs reuse embeddings. (academy.talki-app.fr)
Create a retriever and a retrieval chain that feeds the retrieved chunks to your chat LLM:
from langchain_chroma import Chroma
from langchain_ollama import ChatOllama
from langchain.chains import create_retrieval_chain, create_stuff_documents_chain
# LLM using Ollama (chat model you pulled)
llm = ChatOllama(model="llama3.3", base_url=OLLAMA_BASE)
# Load existing Chroma collection
client = chromadb.PersistentClient(path=CHROMA_DIR)
db = Chroma(client=client, collection_name=COLLECTION, embedding_function=embeddings)
retriever = db.as_retriever(search_kwargs={"k": 4})
combine_chain = create_stuff_documents_chain(llm, prompt=None) # uses default prompt; customize if desired
rag = create_retrieval_chain(retriever, combine_chain)
# Run a query:
resp = rag.invoke({"query": "What are the termination clauses in contract X?"})
print(resp)
LangChain's API has shifted over time (the modern pattern favors create_retrieval_chain / create_stuff_documents_chain or the compatibility package langchain_classic); adapt imports to your installed LangChain version if you see deprecation/compatibility messages. The approach above is the current recommended pattern in recent LangChain guides. (markaicode.com)
You can have a private, zero-cost PDF Q&A system in 2026 with Ollama (local models + embeddings), ChromaDB (persistent local vectors), and LangChain (loaders, splitters, retrieval + chain orchestration). The pattern (load → chunk → embed → store → retrieve → generate) is mature and widely used; once you have Ollama running and a small Python stack installed, building a "chat with your PDFs" app is straightforward and fully offline. For next steps, add a small web UI (Streamlit/Gradio/FastAPI), logging, and automated re‑ingestion when PDFs change.
Selected references and further reading
- Ollama official site and model library (pull models with ollama pull ...). (ollama.com)
- LangChain integrations & reference pages (Ollama & Chroma adapters). (docs.langchain.com)
- ChromaDB docs: persistence and client usage. (docs.trychroma.com)
- Example tutorials and community walkthroughs that follow the same pattern (ingest code and sample projects). (toolbrain.net)
If you want, I can:
- produce a copy‑pasteable repo with working files (ingest.py, query.py, requirements.txt), or
- tailor the example for a specific chat model you prefer (llama3, gemma4, Mistral) and for a particular OS (macOS/Linux/Windows) with exact install commands. Which would you like?