Repolister

A local RAG pipeline for interrogating GitHub repositories

Repolister screenshot

Repolister is a Python application that lets you clone a GitHub repository, chunk it at function level, embed it locally, and query it in plain English — or Japanese, or Chinese, depending on how the prompt is configured. It's built on top of the Ragtime document pipeline, sharing the same PostgreSQL database, embedding model, and Ollama LLM.

Python 3.9+ PostgreSQL + pgvector Ollama / qwen2.5:7b Docker Compose NVIDIA GPU

How it works

Repolister loads a repository through a bronze/silver/gold pipeline:

Bronze

Raw source files, keyed by repo URL and file path.

Silver

Function-level chunks, with roxygen2 docs attached.

Gold

Embeddings via intfloat/multilingual-e5-large, stored in pgvector.

At query time, Repolister embeds your question, retrieves the most relevant function chunks, and asks a local LLM to synthesise an answer — citing the source files it drew from.

Example session

Loading intfloat/multilingual-e5-large on cuda...
  ✓ Ready — querying https://github.com/UchidaMizuki/jpstat

Q: how do I filter data by prefecture?

  Retrieved 6 chunks (closest: 0.1664)

A: To filter data by prefecture, use the activate() function to select
   the area key, then filter() to specify the prefecture name or code.
   For example, to filter for Tokyo and Osaka (R/estat.R):

   census |>
     activate(area) |>
     rekey("pref") |>
     filter(name %in% c("東京都", "大阪府")) |>
     select(code, name)

Currently, Repolister supports R (.r) and R Markdown (.rmd) repositories out of the box, chunked per function or per code fence / prose section. Adding another language means extending the chunker and the supported extensions list.

Full setup instructions (Docker and manual), configuration options, and troubleshooting notes are in the README on GitHub.