Answer questions from your own docs (RAG)
Embed a folder of Markdown docs with the Bike4Mind embeddings API, retrieve the closest sections for a question, and have a model answer only from them, with citations and a log of the questions your docs could not answer.
- Verified
- Every example run against the live API on
- Endpoints
POST /api/v1/embeddings,POST /api/ai/v1/completions- Key scopes
ai:generate- Model
claude-haiku-4-5-20251001- Written in
- Python, curl
Retrieval-augmented generation in one file of standard-library Python. The script splits your docs into sections, embeds them, finds the sections closest to a question, and asks a model to answer from those sections alone. Three properties make it useful for technical documentation rather than a demo:
- It stays current cheaply. Each section is cached by a hash of its text, so after an edit only the changed sections are embedded again.
- It cites. Every answer names the
file#sectionit came from, so a reader can check it. - It records what it could not answer. Questions the docs do not cover are appended to
gaps.jsonl. That file is your list of docs to write next.
You need Python 3.10 or later. No packages to install.
Create an API key
In the Bike4Mind app, open your profile, go to the API Keys tab, and click Create API Key with the AI Generate scope. Embeddings need it; the completions call accepts it too.
export B4M_API_KEY="b4m_live_<your key>"
Look at one embedding
POST /api/v1/embeddings follows the OpenAI embeddings shape. dimensions is optional and shortens
the vector on models that support it.
curl -s -X POST "https://app.bike4mind.com/api/v1/embeddings" \
-H "Authorization: Bearer $B4M_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"text-embedding-3-small","input":["How do I roll back production?"],"dimensions":256}'
The response, with the vector cut to its first three of 256 numbers:
{
"object": "list",
"data": [
{ "object": "embedding", "index": 0, "embedding": [0.06318536150843186, 0.0796208601666945, 0.12564025640982984] }
],
"model": "text-embedding-3-small",
"usage": { "prompt_tokens": 7, "total_tokens": 7 }
}
The response headers carry your quota: x-ratelimit-remaining-minute and
x-ratelimit-remaining-day.
Some docs to search
Any folder of .md files works. To follow along exactly, create three small ones:
mkdir -p docs
cat > docs/deploy.md <<'MD'
# Deploying
Merges to `main` deploy to staging automatically. Production deploys happen only through a
"promote main to prod" pull request, merged as a merge commit.
## Rollback
To roll back production, revert the promotion pull request and merge the revert. Staging
redeploys on its own from `main`.
MD
cat > docs/auth.md <<'MD'
# Authentication
API requests authenticate with an API key sent as `Authorization: Bearer <key>`. Keys are
created in the app under Settings, and each key carries scopes that limit what it can call.
## Rotating keys
Rotate a key by creating a new one, switching your services to it, and then revoking the old key.
MD
cat > docs/limits.md <<'MD'
# Rate limits
Each API key has a per-minute and a per-day request quota. Responses carry
`X-RateLimit-Remaining-Minute` and `X-RateLimit-Remaining-Day` headers. A request over the limit
returns HTTP 429 with a `Retry-After` header.
MD
The script
Save this as docs_qa.py. index() embeds whatever is new, answer() retrieves the top three
sections by cosine similarity and asks the model, and the completions stream is read the same way as
in any other call: join the text of the content events.
"""Answer questions about a folder of Markdown docs, grounded in the docs themselves.
Usage: python3 docs_qa.py ./docs "How do I roll back production?"
"""
import hashlib
import json
import math
import os
import pathlib
import sys
import urllib.request
BASE = "https://app.bike4mind.com"
EMBED_MODEL = "text-embedding-3-small"
CHAT_MODEL = "claude-haiku-4-5-20251001"
CACHE = pathlib.Path(".embeddings.json")
GAPS = pathlib.Path("gaps.jsonl")
def post(path: str, body: dict) -> bytes:
req = urllib.request.Request(
BASE + path,
data=json.dumps(body).encode(),
headers={
"Authorization": f"Bearer {os.environ['B4M_API_KEY']}",
"Content-Type": "application/json",
},
)
with urllib.request.urlopen(req) as res:
return res.read()
def embed(texts: list[str]) -> list[list[float]]:
out = json.loads(post("/api/v1/embeddings", {"model": EMBED_MODEL, "input": texts}))
return [d["embedding"] for d in sorted(out["data"], key=lambda d: d["index"])]
def chunks(folder: pathlib.Path) -> list[dict]:
"""Split every Markdown file on '## ' headings so each chunk is one section."""
result = []
for path in sorted(folder.glob("**/*.md")):
for i, section in enumerate(path.read_text().split("\n## ")):
text = section.strip()
if text:
result.append({"source": f"{path.name}#{i}", "text": text})
return result
def index(folder: pathlib.Path) -> list[dict]:
"""Embed only chunks whose text changed since the last run, so edits stay cheap."""
cache = json.loads(CACHE.read_text()) if CACHE.exists() else {}
items = chunks(folder)
for item in items:
item["hash"] = hashlib.sha256(item["text"].encode()).hexdigest()
todo = [item for item in items if item["hash"] not in cache]
if todo:
for item, vector in zip(todo, embed([item["text"] for item in todo])):
cache[item["hash"]] = vector
print(f"embedded {len(todo)} new or changed chunks", file=sys.stderr)
CACHE.write_text(json.dumps({item["hash"]: cache[item["hash"]] for item in items}))
for item in items:
item["vector"] = cache[item["hash"]]
return items
def cosine(a: list[float], b: list[float]) -> float:
dot = sum(x * y for x, y in zip(a, b))
return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))
def answer(question: str, items: list[dict], k: int = 3) -> str:
q = embed([question])[0]
top = sorted(items, key=lambda item: cosine(q, item["vector"]), reverse=True)[:k]
context = "\n\n".join(f"[{item['source']}]\n{item['text']}" for item in top)
stream = post(
"/api/ai/v1/completions",
{
"model": CHAT_MODEL,
"messages": [
{
"role": "system",
"content": "Answer only from the provided docs and cite the [source] you used. "
"If the docs do not answer the question, reply exactly: NOT IN DOCS",
},
{"role": "user", "content": f"Docs:\n{context}\n\nQuestion: {question}"},
],
},
)
text = ""
for line in stream.decode().splitlines():
if line.startswith("data: {"):
event = json.loads(line[6:])
if event["type"] == "error":
raise RuntimeError(event["message"])
text += event.get("text", "")
if "NOT IN DOCS" in text:
with GAPS.open("a") as f:
f.write(json.dumps({"question": question}) + "\n")
return text
if __name__ == "__main__":
folder, question = pathlib.Path(sys.argv[1]), sys.argv[2]
print(answer(question, index(folder)))
Ask it something the docs cover
python3 docs_qa.py ./docs "How do I roll back production?"
embedded 5 new or changed chunks
According to the docs, to roll back production, you need to [revert the promotion pull request and merge the revert](deploy.md#1). Staging will redeploy on its own from `main`.
[Source: deploy.md#1]
The first run embedded all five sections in one request. The answer cites deploy.md#1, the
"Rollback" section.
Ask it something they do not
python3 docs_qa.py ./docs "Which regions do you support?"
cat gaps.jsonl
NOT IN DOCS
The provided documentation covers deployment procedures and key rotation practices, but does not contain information about which regions are supported.
{"question": "Which regions do you support?"}
No embedded line this time: nothing changed, so the second run made two API calls (one embedding
for the question, one completion) instead of three. The model declined to guess, and the question
landed in gaps.jsonl.
Where to take it
- Bigger corpora. The script holds every vector in memory and compares them in plain Python, which is fine for hundreds of sections. Past that, put the vectors in a vector index. One request accepts up to 2048 inputs, and the input count is further capped so inputs times dimensions stays within 196,608 values, so page large first runs.
- Chunking. Splitting on
##headings keeps each chunk a coherent section. If your docs have long sections, split those further, and keep the heading in each piece. - Same model, same space. Query and documents must be embedded with the same model and
dimensions. If you change either, delete.embeddings.jsonand rebuild. - Managed retrieval. Bike4Mind data lakes index and search files on the server instead. That API needs the data-lake scopes and was not run for this guide, so it is not shown here.
Something here no longer matches what the API returns? Tell us at support@bike4mind.com and we will rerun it. Full endpoint reference: API explorer.