AI Assistant
A self-hosted AI workbench that answers questions and writes code from a private knowledge base. Retrieval-augmented generation end to end, with every model running on local hardware.
Python, FastAPI, OpenSearch, sentence-transformers, Hugging Face Transformers, llama-cpp-python, Next.js, TypeScript, shadcn/ui, WebSockets, Docker

The problem
Some products run where there is no internet access, or under security rules strict enough that sending documentation to a hosted AI service is off the table. The teams behind them still need answers from their own documentation, and general-purpose models don't know their internal tools or conventions.
Approach
Retrieval-augmented generation, end to end. Documents go through a custom parser that keeps their structure, are split into chunks, embedded and indexed in OpenSearch. At question time the assistant retrieves the most relevant passages and hands them to the model with the question, so answers rest on the organisation's own material and show their sources.
Retrieval is configurable rather than fixed: semantic, keyword or hybrid search, optional re-ranking, result limits, a minimum relevance score and filters by file type and category, each with a sensible default in the settings.


Generation runs through reusable prompt templates. The assistant assembles the retrieved documentation and matching examples into the template, runs the active model and keeps every exchange in a searchable history with its prompt, sources and metadata, exportable as JSON or CSV.


Models are managed inside the app: a registry of curated and Hugging Face models, a download queue with live progress, and one device manager that places each model on an NVIDIA GPU, Apple Silicon or the CPU. Document summarisation and code-similarity analysis run on the same local stack.

Built for air-gapped environments
Everything runs on the machine it is installed on: the models, the embeddings, the search index and the history. Once models are downloaded, the assistant needs no outside connection, which makes the approach usable where hosted AI is not an option.
What was hard
Local models have far less room for context than hosted ones, so retrieval has to be precise. Deciding how much to retrieve, how to rank it and how to fit it into the prompt took most of the iteration.
Keeping the interface responsive while multi-gigabyte models download and load meant moving that work off the request path and streaming status to the browser over WebSockets. Running the same code on different GPUs and on CPU-only machines needed a single place that decides where each model runs.
Outcome
A working assistant covering the full RAG loop, from ingestion to answers grounded in their sources, with model management, summarisation and code-similarity analysis around it, all on local hardware. Richer document parsing is in progress.