Answer questions from your own docs (RAG)

Embed a folder of Markdown docs with the Bike4Mind embeddings API, retrieve the closest sections for a question, and have a model answer only from them, with citations and a log of the questions your docs could not answer.

Verified
Every example run against the live API on
Endpoints
POST /api/v1/embeddings, POST /api/ai/v1/completions
Key scopes
ai:generate
Model
claude-haiku-4-5-20251001
Written in
Python, curl

Retrieval-augmented generation in one file of standard-library Python. The script splits your docs into sections, embeds them, finds the sections closest to a question, and asks a model to answer from those sections alone. Three properties make it useful for technical documentation rather than a demo:

  • It stays current cheaply. Each section is cached by a hash of its text, so after an edit only the changed sections are embedded again.
  • It cites. Every answer names the file#section it came from, so a reader can check it.
  • It records what it could not answer. Questions the docs do not cover are appended to gaps.jsonl. That file is your list of docs to write next.

You need Python 3.10 or later. No packages to install.

Create an API key

In the Bike4Mind app, open your profile, go to the API Keys tab, and click Create API Key with the AI Generate scope. Embeddings need it; the completions call accepts it too.

export B4M_API_KEY="b4m_live_<your key>"

Look at one embedding

POST /api/v1/embeddings follows the OpenAI embeddings shape. dimensions is optional and shortens the vector on models that support it.

curl -s -X POST "https://app.bike4mind.com/api/v1/embeddings" \
  -H "Authorization: Bearer $B4M_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"text-embedding-3-small","input":["How do I roll back production?"],"dimensions":256}'

The response, with the vector cut to its first three of 256 numbers:

{
  "object": "list",
  "data": [
    { "object": "embedding", "index": 0, "embedding": [0.06318536150843186, 0.0796208601666945, 0.12564025640982984] }
  ],
  "model": "text-embedding-3-small",
  "usage": { "prompt_tokens": 7, "total_tokens": 7 }
}

The response headers carry your quota: x-ratelimit-remaining-minute and x-ratelimit-remaining-day.

Any folder of .md files works. To follow along exactly, create three small ones:

mkdir -p docs
cat > docs/deploy.md <<'MD'
# Deploying

Merges to `main` deploy to staging automatically. Production deploys happen only through a
"promote main to prod" pull request, merged as a merge commit.

## Rollback

To roll back production, revert the promotion pull request and merge the revert. Staging
redeploys on its own from `main`.
MD
cat > docs/auth.md <<'MD'
# Authentication

API requests authenticate with an API key sent as `Authorization: Bearer <key>`. Keys are
created in the app under Settings, and each key carries scopes that limit what it can call.

## Rotating keys

Rotate a key by creating a new one, switching your services to it, and then revoking the old key.
MD
cat > docs/limits.md <<'MD'
# Rate limits

Each API key has a per-minute and a per-day request quota. Responses carry
`X-RateLimit-Remaining-Minute` and `X-RateLimit-Remaining-Day` headers. A request over the limit
returns HTTP 429 with a `Retry-After` header.
MD

The script

Save this as docs_qa.py. index() embeds whatever is new, answer() retrieves the top three sections by cosine similarity and asks the model, and the completions stream is read the same way as in any other call: join the text of the content events.

"""Answer questions about a folder of Markdown docs, grounded in the docs themselves.

Usage: python3 docs_qa.py ./docs "How do I roll back production?"
"""

import hashlib
import json
import math
import os
import pathlib
import sys
import urllib.request

BASE = "https://app.bike4mind.com"
EMBED_MODEL = "text-embedding-3-small"
CHAT_MODEL = "claude-haiku-4-5-20251001"
CACHE = pathlib.Path(".embeddings.json")
GAPS = pathlib.Path("gaps.jsonl")


def post(path: str, body: dict) -> bytes:
    req = urllib.request.Request(
        BASE + path,
        data=json.dumps(body).encode(),
        headers={
            "Authorization": f"Bearer {os.environ['B4M_API_KEY']}",
            "Content-Type": "application/json",
        },
    )
    with urllib.request.urlopen(req) as res:
        return res.read()


def embed(texts: list[str]) -> list[list[float]]:
    out = json.loads(post("/api/v1/embeddings", {"model": EMBED_MODEL, "input": texts}))
    return [d["embedding"] for d in sorted(out["data"], key=lambda d: d["index"])]


def chunks(folder: pathlib.Path) -> list[dict]:
    """Split every Markdown file on '## ' headings so each chunk is one section."""
    result = []
    for path in sorted(folder.glob("**/*.md")):
        for i, section in enumerate(path.read_text().split("\n## ")):
            text = section.strip()
            if text:
                result.append({"source": f"{path.name}#{i}", "text": text})
    return result


def index(folder: pathlib.Path) -> list[dict]:
    """Embed only chunks whose text changed since the last run, so edits stay cheap."""
    cache = json.loads(CACHE.read_text()) if CACHE.exists() else {}
    items = chunks(folder)
    for item in items:
        item["hash"] = hashlib.sha256(item["text"].encode()).hexdigest()
    todo = [item for item in items if item["hash"] not in cache]
    if todo:
        for item, vector in zip(todo, embed([item["text"] for item in todo])):
            cache[item["hash"]] = vector
        print(f"embedded {len(todo)} new or changed chunks", file=sys.stderr)
    CACHE.write_text(json.dumps({item["hash"]: cache[item["hash"]] for item in items}))
    for item in items:
        item["vector"] = cache[item["hash"]]
    return items


def cosine(a: list[float], b: list[float]) -> float:
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))


def answer(question: str, items: list[dict], k: int = 3) -> str:
    q = embed([question])[0]
    top = sorted(items, key=lambda item: cosine(q, item["vector"]), reverse=True)[:k]
    context = "\n\n".join(f"[{item['source']}]\n{item['text']}" for item in top)
    stream = post(
        "/api/ai/v1/completions",
        {
            "model": CHAT_MODEL,
            "messages": [
                {
                    "role": "system",
                    "content": "Answer only from the provided docs and cite the [source] you used. "
                    "If the docs do not answer the question, reply exactly: NOT IN DOCS",
                },
                {"role": "user", "content": f"Docs:\n{context}\n\nQuestion: {question}"},
            ],
        },
    )
    text = ""
    for line in stream.decode().splitlines():
        if line.startswith("data: {"):
            event = json.loads(line[6:])
            if event["type"] == "error":
                raise RuntimeError(event["message"])
            text += event.get("text", "")
    if "NOT IN DOCS" in text:
        with GAPS.open("a") as f:
            f.write(json.dumps({"question": question}) + "\n")
    return text


if __name__ == "__main__":
    folder, question = pathlib.Path(sys.argv[1]), sys.argv[2]
    print(answer(question, index(folder)))

Ask it something the docs cover

python3 docs_qa.py ./docs "How do I roll back production?"
embedded 5 new or changed chunks
According to the docs, to roll back production, you need to [revert the promotion pull request and merge the revert](deploy.md#1). Staging will redeploy on its own from `main`.

[Source: deploy.md#1]

The first run embedded all five sections in one request. The answer cites deploy.md#1, the "Rollback" section.

Ask it something they do not

python3 docs_qa.py ./docs "Which regions do you support?"
cat gaps.jsonl
NOT IN DOCS

The provided documentation covers deployment procedures and key rotation practices, but does not contain information about which regions are supported.
{"question": "Which regions do you support?"}

No embedded line this time: nothing changed, so the second run made two API calls (one embedding for the question, one completion) instead of three. The model declined to guess, and the question landed in gaps.jsonl.

Where to take it

  • Bigger corpora. The script holds every vector in memory and compares them in plain Python, which is fine for hundreds of sections. Past that, put the vectors in a vector index. One request accepts up to 2048 inputs, and the input count is further capped so inputs times dimensions stays within 196,608 values, so page large first runs.
  • Chunking. Splitting on ## headings keeps each chunk a coherent section. If your docs have long sections, split those further, and keep the heading in each piece.
  • Same model, same space. Query and documents must be embedded with the same model and dimensions. If you change either, delete .embeddings.json and rebuild.
  • Managed retrieval. Bike4Mind data lakes index and search files on the server instead. That API needs the data-lake scopes and was not run for this guide, so it is not shown here.

Something here no longer matches what the API returns? Tell us at support@bike4mind.com and we will rerun it. Full endpoint reference: API explorer.