jamwithai/production-agentic-rag-course: Open-Source Production Agentic RAG Engineering Pipeline

An open-source reference course and production architecture covering OpenSearch 2.19 hybrid search, Docling paper parsing, Airflow pipelines, LangGraph agentic

tau · October 3, 2026

#RAG #AgenticAI #OpenSearch #LangGraph #FastAPI #Python

jamwithai/production-agentic-rag-course: Open-Source Production Agentic RAG Engineering Pipeline

Moving beyond rudimentary proof-of-concept RAG implementations, jamwithai/production-agentic-rag-course is an open-source project providing a comprehensive, production-ready retrieval and agentic orchestration architecture. Structured as a hands-on 7-week curriculum centered around an automated arXiv paper curation system (arxiv-paper-curator), it walks developers through every stage of the modern enterprise RAG lifecycle—from ingestion and hybrid retrieval to grading, guardrails, and observability.

Production agentic RAG architecture diagram and engineering workflow with OpenSearch and LangGraph

Image source: jamwithai (GitHub)

In contemporary AI application development, basic vector database similarity search often struggles with vocabulary mismatch, dense domain jargon, and ungrounded hallucinations. The emerging industry standard combines keyword and vector retrieval with dynamic state-machine agents. This repository translates these production requirements into a clean, fully containerized codebase.

OpenSearch 2.19 Hybrid Search and Ingestion Pipelines

The core retrieval layer is powered by OpenSearch 2.19, integrating both lexical and semantic search modalities.

  • Hybrid Indexing: Combines traditional BM25 keyword matching with dense vector embeddings to achieve high recall and precision across specialized terminology and conversational queries alike.
  • Document Parsing: Utilizes IBM's Docling parser to extract structured text, tables, and mathematical formulas from complex academic PDF papers.
  • Orchestration and Metadata: Employs Apache Airflow to automate the continuous ingestion of arXiv papers, backed by PostgreSQL 16 for structured metadata tracking and chunk index management.

Developers can inspect indexes and test queries interactively via the OpenSearch REST API at http://localhost:9200 and the OpenSearch Dashboards web UI at http://localhost:5601.

LangGraph-Driven Agentic Routing and Guardrails

Rather than statically feeding retrieved contexts directly into generation prompts, the system uses LangGraph to manage an adaptive decision-making graph.

  • Document Grading: An LLM-based grading node evaluates whether retrieved document chunks are relevant to the user query, filtering out irrelevant context.
  • Adaptive Query Rewriting: If the initial retrieval fails to return sufficient relevant evidence, the agent analyzes the gap and rewrites the query for a subsequent search iteration.
  • Guardrail Node: Intercepts queries prior to execution, conducting input validation and domain boundary detection to catch out-of-scope or adversarial requests.
  • Interactive Serving: Exposes the agent workflow through a Telegram Bot interface for real-time conversational testing.

Modern Python 3.12, UV, and FastAPI Infrastructure

The repository is built on modern Python tooling and standardized container orchestration for rapid local onboarding.

  • Backend & Tooling: Uses Python 3.12+ managed via the UV package manager, exposing async REST endpoints through FastAPI.
  • Local Inference & Caching: Supports local LLM serving with Ollama and reduces repetitive latency via a Redis caching layer.
  • Observability: Integrates Langfuse tracing out of the box to track per-step latency, token consumption, and retrieval evaluation metrics.

System Requirements and Operational Notes

  • Resource Requirements: Running the full multi-service Docker Compose stack (OpenSearch, PostgreSQL, Airflow, Redis, FastAPI) requires Docker Desktop, at least 8GB of RAM, and 20GB of available disk space.
  • API Credentials: While local Ollama models can be used, utilizing cloud embeddings or remote tracing requires setting up Jina Embeddings API keys and Langfuse account credentials in the environment configuration.

Sources