Open-source, zero-cost AI site search & chatbot widget. 100% Cloudflare native.
Drop one <script> tag on your docs site and get a streaming RAG chatbot that answers from your content, with sources.
No OpenAI key. No Pinecone. No servers. No bill.
Screenshots
| Chat widget embedded in a docs site | Admin UI (/admin) |
|---|---|
![]() |
![]() |
Streaming answers with source chips on the left; the built-in admin page for registering sitemaps, driving indexing and reviewing what users ask on the right.
Why
Hosted "chat with your docs" widgets (Chatbase, CustomGPT, Mendable, ...) charge $20-$400/month for something that is, under the hood, a crawler, an embedding model, a vector store and an LLM. Cloudflare now ships every one of those pieces with a generous free tier:
| Piece | Cloudflare service | Free tier (at time of writing) |
|---|---|---|
| API / crawler | Workers + Hono | 100k requests/day |
| Embeddings + LLM | Workers AI (bge-small-en-v1.5, bge-reranker-base, llama-3.1-8b-instruct-fp8) |
10k neurons/day |
| Vector search | Vectorize (384-dim, cosine) | 5M stored dims, 30M queried dims/month |
| Metadata + chat log | D1 (SQLite) | 5 GB, 5M reads/day |
| Background indexing | Cron Triggers | included |
| Widget hosting | Static Assets | included |
DocFlare AI wires them together into a single Worker you deploy with one command.
Features
- One-tag embed:
<script src=".../widget.js" data-site-id="my-docs">. Vanilla JS, Shadow DOM, < 10 KB, zero dependencies, no CSS leaks. - Streaming answers over Server-Sent Events with source chips, in the language of the question.
- Sitemap crawler: sitemap index support, gzip sitemaps, boilerplate stripping (nav/header/footer/scripts) via the native
HTMLRewriter. - Section-aware chunking: headings start new chunks, so a landing page's "Open source" card is not buried in a chunk about pricing. Chunks are ~350 tokens with 40-token overlap, sized with a per-script token estimate so Vietnamese/CJK pages stay inside the embedding model's 512-token window.
- Two-stage retrieval: the 20 best vector matches are re-scored by the
bge-reranker-basecross-encoder before the top-K go to the LLM. Questions that merely share vocabulary with the wrong page (the site name on legal pages, a product name) land on the paragraph that actually answers them. - Multi-site / multi-tenant: one deployment can serve many documentation sites, each isolated in its own Vectorize namespace and locked to its own domain.
- Free-tier safe by design: indexing runs in small resumable batches (cron + on-demand), the LLM is skipped when nothing relevant is retrieved, and AI rate limits degrade gracefully.
- Dark / light / auto theme, custom accent colour, left/right position, keyboard friendly.
- Chat history in D1 so you can see what users ask and where your docs have gaps.
- Built-in admin UI at
/admin: register sitemaps, watch indexing progress, process batches, re-crawl, delete sites, copy the embed snippet and browse recent questions. Single static HTML file, no build step. - TypeScript end-to-end, Hono router, strict types, no
any. Bun for tooling and tests.
How it works
flowchart LR
subgraph Host["Your docs site"]
W["widget.js<br/>(Shadow DOM)"]
end
subgraph CF["Cloudflare (free tier)"]
direction TB
API["Worker + Hono<br/>src/index.ts"]
AI["Workers AI<br/>bge-small-en-v1.5<br/>bge-reranker-base<br/>llama-3.1-8b-instruct-fp8"]
VX["Vectorize<br/>384-dim cosine<br/>namespace = siteId"]
D1["D1<br/>sites / pages / queries"]
CRON["Cron trigger<br/>every minute"]
ASSETS["Static assets<br/>public/"]
end
SM["sitemap.xml + HTML pages"]
W -- "POST /api/chat (SSE)" --> API
ASSETS -- "GET /widget.js" --> W
API -- "embed question" --> AI
API -- "topK query" --> VX
API -- "RAG prompt" --> AI
API -- "log query" --> D1
CRON -- "process pending pages" --> API
API -- "fetch + extract" --> SM
API -- "embed chunks" --> AI
API -- "upsert vectors" --> VX
Indexing: POST /api/sites/index reads the sitemap and queues every HTML URL in D1. A cron trigger (or repeated calls to /process) then crawls a few pages per invocation: extract text, chunk, embed, upsert to Vectorize, mark the page indexed.
Chat: the widget sends the question; the Worker embeds it, pulls the 20 closest chunks from the site's namespace, reranks them with a cross-encoder, keeps the top-K (max 2 per page), builds a grounded prompt, streams Llama 3.1's answer back as SSE and logs the exchange in D1.
See docs/ARCHITECTURE.md for the full design and docs/API_REFERENCE.md for the endpoints.
Quick start
Prerequisites
- A free Cloudflare account with Workers enabled.
- Bun 1.1+ (used for installing, scripts, building the widget and running tests).
1. Clone and install
git clone https://github.com/p10node/docflare-ai.git
cd docflare-ai
bun install
bunx wrangler login2. Create the D1 database
bunx wrangler d1 create docflare-db
Copy the database_id from the output. Either paste it into wrangler.toml:
[[d1_databases]] binding = "DB" database_name = "docflare-db" database_id = "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx" # <- paste here
…or, to keep it out of git, put it in a .env file instead (useful when you push your fork):
echo 'D1_DATABASE_ID=xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' > .env
With .env set, bun run deploy and bun run db:migrate generate wrangler.local.toml (gitignored) from wrangler.toml with the id filled in and use that. wrangler.toml keeps the placeholder.
Apply the schema:
3. Create the Vectorize index
bunx wrangler vectorize create docflare-index --dimensions=384 --metric=cosine
(384 dimensions matches @cf/baai/bge-small-en-v1.5. The index name is already set in wrangler.toml.)
4. Set the admin token
The management endpoints are protected by a bearer token. Pick something long and random:
bunx wrangler secret put ADMIN_TOKEN
5. Deploy
bun run deploy builds the widget and deploys (the schema was applied in step 2; re-run bun run db:migrate whenever migrations/ gains a new file). Wrangler prints your Worker URL, e.g. https://docflare-ai.<your-subdomain>.workers.dev. Open it to see the demo page with the widget mounted.
Deploy to Cloudflare button: the button at the top of this README clones the repo into your GitHub account, provisions the Worker, D1 database and Vectorize index from
wrangler.toml, asks forADMIN_TOKEN(from.dev.vars.example) and runsbun run deploy. Two values the wizard cannot read fromwrangler.toml:
- Vectorize index: enter Dimensions = 384 and Metric = cosine (the form leaves them blank; any other shape breaks indexing).
ADMIN_TOKEN: replace the example value with something long and random, e.g.openssl rand -hex 32.The build does not apply the D1 schema (the build token has no D1 access), so run the migration once from your new repo after the first deploy. Cloudflare writes your
database_idinto that repo'swrangler.tomlwhile provisioning, so no.envis needed (if it still shows the placeholder, do step 2):git clone https://github.com/<you>/docflare-ai.git && cd docflare-ai bun install && bunx wrangler login bun run db:migrate bunx wrangler vectorize get docflare-index # expect dimensions: 384, metric: cosineIf the index is missing or has a different shape, run step 3 (delete it first with
bunx wrangler vectorize delete docflare-index) and redeploy. If the wizard did not ask forADMIN_TOKEN, run step 4.
Index your documentation
export WORKER_URL="https://docflare-ai.<your-subdomain>.workers.dev" export ADMIN_TOKEN="the-token-you-set" curl -X POST "$WORKER_URL/api/sites/index" \ -H "Authorization: Bearer $ADMIN_TOKEN" \ -H "Content-Type: application/json" \ -d '{"siteId":"my-docs","sitemapUrl":"https://docs.example.com/sitemap.xml"}'
Response:
{
"siteId": "my-docs",
"domain": "docs.example.com",
"sitemapUrl": "https://docs.example.com/sitemap.xml",
"queued": 142,
"progress": { "siteId": "my-docs", "total": 142, "pending": 142, "indexed": 0, "failed": 0 },
"next": "Pending pages are crawled 3 at a time by the cron trigger. Call POST /api/sites/my-docs/process repeatedly to finish faster."
}Prefer a UI? Open https://docflare-ai.<your-subdomain>.workers.dev/admin, paste the admin token and use the Register / re-index a site form. The same page shows progress bars per site, an Index all pending button that loops over /process until the queue is empty, and the recent questions log.
Pages are now crawled automatically, a few per minute, by the cron trigger. To finish faster, drive the queue yourself from the CLI:
until curl -sf -X POST "$WORKER_URL/api/sites/my-docs/process" \ -H "Authorization: Bearer $ADMIN_TOKEN" | grep -q '"pending":0'; do sleep 1 done
Check progress at any time:
curl "$WORKER_URL/api/sites/my-docs/status" -H "Authorization: Bearer $ADMIN_TOKEN"
Re-running /api/sites/index for the same siteId re-queues every page, so a cron job or CI step can keep the index fresh. Vectorize is eventually consistent: freshly upserted chunks become searchable within a few seconds.
After upgrading DocFlare to a version that changes chunking (see the changelog in commit messages), re-index every site once (Re-crawl sitemap in the admin UI or
POST /api/sites/index) so pages are re-chunked. Old vectors keep working until then.
Embed the widget
Add one line before </body> on any page of the registered domain:
<script src="https://docflare-ai.<your-subdomain>.workers.dev/widget.js" data-site-id="my-docs" defer></script>
Widget options
| Attribute | Default | Description |
|---|---|---|
data-site-id |
required | The siteId you indexed. |
data-api |
script origin | Base URL of the Worker, if you serve widget.js from elsewhere (e.g. a CDN). |
data-theme |
auto |
light, dark or auto (follows prefers-color-scheme). |
data-color |
#f6821f |
Accent colour for the bubble and buttons. |
data-position |
right |
right or left corner. |
data-title |
Ask AI |
Panel header text. |
data-placeholder |
Ask a question about these docs… |
Input placeholder. |
data-welcome |
Hi! Ask me anything… |
First assistant message. |
A tiny JavaScript API is exposed as window.DocFlare:
DocFlare.open(); // open the panel DocFlare.close(); DocFlare.ask('How do I deploy?'); // open and submit a question
Only pages served from the site's registered domain (and anything in ALLOWED_ORIGINS) can call /api/chat for that siteId. Requests from other origins get a 403.
Configuration
Non-secret settings live in the [vars] section of wrangler.toml:
| Variable | Default | Purpose |
|---|---|---|
MAX_PAGES |
200 |
Max URLs taken from a sitemap per indexing run. |
INDEX_BATCH_SIZE |
3 |
Pages crawled + embedded per cron tick or /process call. Raise to 10-20 on the Workers Paid plan. |
TOP_K |
4 |
Chunks handed to the LLM per question (max 20). |
RERANK |
true |
Rerank the 20 best vector matches with @cf/baai/bge-reranker-base before picking TOP_K. false = vector order only. |
ALLOWED_ORIGINS |
"" |
Comma-separated extra origins allowed to call /api/chat for any site. * allows everything. |
Secrets (set with bunx wrangler secret put NAME):
| Secret | Purpose |
|---|---|
ADMIN_TOKEN |
Bearer token for /api/sites/*. Required; management routes return 503 until it is set. |
Want a bigger model or a different embedding size? Change CHAT_MODEL / EMBEDDING_MODEL in src/services/ai.ts and create a Vectorize index with the matching dimensions.
Local development
cp .dev.vars.example .dev.vars # sets ADMIN_TOKEN for local runs bunx wrangler d1 migrations apply docflare-db --local bun run dev # http://localhost:8787
Notes:
- Workers AI and Vectorize have no local emulator;
wrangler devproxies those bindings to your Cloudflare account (you must be logged in). D1 runs locally. wrangler devuses a local D1 and ignoresdatabase_id, so neither the real id inwrangler.tomlnor.envis needed for local runs. Onlybun run deployandbun run db:migrateneed it.- The Vectorize index must exist remotely before
wrangler devcan use it. - Test the cron handler locally with
bun run dev:cron, thencurl "http://localhost:8787/__scheduled?cron=*+*+*+*+*". bun testruns the unit tests (token estimate, section chunker, HTML extraction, sitemap filtering, CORS policy, reranker, prompt builder).bun run build:checktype-checks and performs a dry-run bundle without deploying.bun run build:widgetminifieswidget/widget.src.jsintopublic/widget.jswithbun build(run after editing the widget).
API overview
| Method | Path | Auth | Description |
|---|---|---|---|
GET |
/api/health |
none | Liveness + model names. |
POST |
/api/sites/index |
admin | Register a site and queue its sitemap URLs. |
POST |
/api/sites/:siteId/process |
admin | Crawl and embed the next batch of pending pages. |
GET |
/api/sites/:siteId/status |
admin | Indexing progress. |
DELETE |
/api/sites/:siteId |
admin | Remove a site, its pages, chat log and vectors. |
POST |
/api/chat |
origin check | Ask a question (JSON or SSE stream). |
GET |
/widget.js |
none | Embeddable widget. |
Full request/response shapes, the SSE protocol and an OpenAPI document are in docs/API_REFERENCE.md.
Staying inside the free tier
Rough budget for a typical documentation site, based on Cloudflare pricing as of 2026-10-01 (check the Workers AI, Vectorize, D1 and Workers pricing pages for up-to-date numbers):
- Indexing:
bge-smallcosts 1,841 neurons per million input tokens, i.e. roughly 1 neuron per page. A 1,000-page site fits comfortably in a day's free quota. - Chat: one answered question costs roughly 30-70 neurons: the prompt is 1.5k-3.5k tokens (system prompt +
TOP_K=4chunks of ~350 tokens + conversation history) at 13,778 neurons/M, the answer is 300-768 tokens at 26,128 neurons/M, the query embedding is negligible and reranking 20 candidates (~6k tokens at 283 neurons/M) costs about 2 neurons. With 10k free neurons/day that is on the order of 120-280 answered questions per day for free. Questions with no relevant context never hit the LLM and cost ~0. - Vectorize: 5M free stored dimensions / 384 ≈ 13,000 chunks ≈ 1,000-2,500 pages across all sites (section-aware chunks are smaller and more numerous than fixed 500-token blocks). Each query reads 384 dimensions, so even 2,500 questions/day stay under the 30M queried dimensions/month allowance.
- D1 and Workers requests: one chat is 1 row read + 1 row written; the per-minute cron adds ~1.4k reads/day. Both are orders of magnitude below the free limits.
- CPU time: the Workers Free plan allows 10 ms CPU per invocation. Crawling is I/O bound and
HTMLRewriteris native, so a batch of 3 pages stays under the limit; on the Paid plan raiseINDEX_BATCH_SIZE.
Beyond the free tier
Workers AI is the only piece you will outgrow. Past the daily neuron allowance the AI binding fails and /api/chat returns 429 rate_limited until the quota resets. To go further, move to the Workers Paid plan ($5/month): it keeps the 10k free neurons/day and bills $0.011 per 1,000 neurons beyond that. No code changes are needed.
| Questions / day | Plan | Workers AI | Total / month (approx.) |
|---|---|---|---|
| up to ~280 | Free | $0 | $0 |
| 1,000 | Paid | $9-24 | $14-29 |
| 5,000 | Paid | $56-132 | $61-137 |
That is about $0.0004-0.0009 per question beyond the free allowance. The spread comes from how long the retrieved chunks, the history and the answers are.
To cut the cost per question:
- Lower
TOP_Kfrom 4 to 3 (about 25% fewer prompt tokens, slightly less context for the model). - Lower
MAX_OUTPUT_TOKENSinsrc/services/ai.ts(768 by default); output tokens are the most expensive kind. - Trim
MAX_HISTORY_TURNS/MAX_HISTORY_CHARSinsrc/services/ai.tsif follow-up questions are rare on your site. - Swap
CHAT_MODELfor a cheaper model from the Workers AI catalog; the response parser accepts both Llama-style and OpenAI-style outputs.
Limitations and roadmap
- English-optimised embedding model (
bge-small-en). For multilingual docs switch to@cf/baai/bge-m3(1024 dims) and recreate the index. - JavaScript-rendered sites must expose server-rendered HTML (most docs generators do).
- No authentication for private docs yet (the widget is meant for public sites).
- Planned: incremental re-crawl using sitemap
<lastmod>, feedback thumbs on answers, Markdown/llms.txt ingestion, per-site widget settings in the admin UI.
Project structure
docflare-ai/
├── docs/
│ ├── images/ # README screenshots
│ ├── ARCHITECTURE.md # System design & RAG data flow
│ └── API_REFERENCE.md # REST + SSE endpoint reference, OpenAPI
├── migrations/
│ └── 0000_init.sql # D1 schema
├── scripts/
│ └── wrangler.ts # Wrapper: swaps D1_DATABASE_ID from .env (optional) into a gitignored wrangler.local.toml
├── public/
│ ├── _headers # Static asset headers (CORS / caching for widget.js)
│ ├── admin.html # Admin UI (served at /admin)
│ ├── index.html # Demo page
│ └── widget.js # Minified embeddable widget (built from widget/)
├── widget/
│ └── widget.src.js # Readable widget source (bun build → public/widget.js)
├── test/
│ └── unit.test.ts # bun test: chunker, crawler filters, CORS, prompt
├── src/
│ ├── db/schema.ts # D1 statements, row types, repository
│ ├── services/
│ │ ├── ai.ts # Workers AI embeddings, RAG prompt, generation
│ │ ├── crawler.ts # Sitemap fetcher, HTMLRewriter text extractor, chunker
│ │ ├── indexer.ts # Batch indexing pipeline
│ │ └── vectorize.ts # Vectorize upsert / query / delete
│ ├── utils/
│ │ ├── cors.ts # Multi-domain CORS + origin policy
│ │ └── hash.ts # SHA-256 ids, constant-time compare
│ ├── index.ts # Hono app, routes, cron handler
│ └── types.ts # Env bindings & shared types
├── wrangler.toml
├── .dev.vars.example # ADMIN_TOKEN etc. for `wrangler dev`
├── package.json
├── bun.lock
└── tsconfig.json
Contributing
Issues and pull requests are welcome. Run bun test and bun run build:check before opening a PR; they run the unit tests, type-check and dry-run bundle the Worker.

