Jevbox: An Open-Source Self-Organizing Document Drive Powered by Hierarchical Classification Without Vector DBs

Extend has released Jevbox, an open-source document drive that uses the TypeSafe Jev decision model to navigate folders, documents, and sections hierarchically

tau · October 3, 2026

#Jevbox #Jev #MCP #OpenSource #DocumentSearch #DevTools

Jevbox: An Open-Source Self-Organizing Document Drive Powered by Hierarchical Classification Without Vector DBs

On October 3, 2026, Andrew Luo and Kushal Byatnal of Extend unveiled Jevbox, an open-source document drive project that automatically categorizes files and answers user questions through hierarchical classification powered by the TypeSafe Jev decision model. Built on top of Extend Parse and Jev, the entire project has been released as open source on GitHub (extend-hq/jevbox).

Architecture diagram illustrating Jevbox open-source document drive with hierarchical folder classification and retrieval-time permission controls

Image source: @kushalbyatnal via X

Unlike conventional Retrieval-Augmented Generation (RAG) pipelines that chunk documents into dense high-dimensional vectors and query vector databases for semantic similarity, Jevbox introduces a navigation pattern that mirrors human directory browsing—progressively narrowing down relevant targets through step-by-step classification.

Hierarchical Search and Automated Document Organization Without Vector Databases

The core architecture of Jevbox entirely bypasses embeddings and vector databases across its ingestion and query workflows.

When a user uploads a document, the Jev decision model analyzes its content, categorizes it, and files it directly into the appropriate location within a pre-structured folder tree. When a user asks a natural language question, the search process follows a top-down hierarchical flow rather than computing distance across vector spaces.

  • Hierarchical Search Funnel: Jev first identifies the most relevant top-level folder, selects the candidate document within that folder, and then isolates the exact section containing the answer.
  • Precise Page and Section Citations: Because retrieval targets structured structural segments rather than arbitrary text chunks, every answer generated by Jevbox explicitly cites the source page and section from which the facts were retrieved.

This hierarchical approach avoids the compute overhead and infrastructure management costs of high-dimensional vector databases while mitigating context fragmentation that frequently occurs during arbitrary chunking.

Retrieval-Time Permission Enforcement and Native MCP Server for AI Agents

Jevbox also integrates access control and ecosystem connectivity designed for real-world team environments.

A common challenge in enterprise document retrieval is ensuring that strict access permissions are respected across varying user privilege levels. Jevbox enforces security policies dynamically at retrieval time.

  • Dynamic Permission Checks: If a user's access to a source document is revoked, their access to answers derived from that document is immediately blocked at query time, without requiring vector index re-indexing or batch cache invalidation.
  • Native MCP (Model Context Protocol) Server: Jevbox ships with a built-in MCP server, allowing external AI coding agents and workspace tools to query the document repository. Agents operate under the exact same permission-aware access boundaries as human users.

Deployment Paths and Operational Considerations

Jevbox provides flexible options for both local self-hosting and cloud deployments.

The GitHub repository provides instructions for developers looking to self-host the application on their own infrastructure, alongside a one-click deployment template for the Render cloud platform.

When considering adoption, developers should note key operational requirements. Jevbox relies on the TypeSafe Jev decision and classification model API rather than local vector calculations, meaning a configured Jev API environment is required to operate the workflow. Furthermore, in environments with highly complex unstructured hierarchies or multi-version document trees, the accuracy of document filing and retrieval depends on thoughtful folder tree design and well-defined classification prompts.

Sources