Decision log
Newest first. Each entry: what was decided, what it was weighed against, and the measurement.
2026-09-10 — Postgres self-hosted on the search VPS, not Neon
Section titled “2026-09-10 — Postgres self-hosted on the search VPS, not Neon”Neon’s free tier caps egress at 5 GB/month; the test site used 5.06 GB in its first day, mostly crawlers walking the sitemap and listing pages, and the compute never idled because Hangar’s health probe hits the API every few seconds. Options were Neon’s paid plan (~$19/month), or Postgres in the existing Compose stack on the box we already pay for. Chose the latter: no caps, TLS with a self-signed certificate, port open only to the Hangar IP, pg_stat_statements on so the next surprise is visible. Nightly dumps go to the Hangar box.
2026-09-10 — Copy the search database instead of re-embedding
Section titled “2026-09-10 — Copy the search database instead of re-embedding”Building the semantic indexes means embedding ~1 M sentences and ~360 k passages with bge-m3, about 10 GPU-hours here, then Meilisearch building vector graphs, which ran at ~14 docs/s on the Mac and would take over a day on 2 vCPUs. The local instance already held the finished 25 GB database. Copying it took most of a day only because of a bad link and three script bugs, but it needs no GPU and no graph build; the transfer itself is ~35 minutes at 100 Mbit/s. Same Meilisearch version on both ends is mandatory.
2026-09-09 — Search box on its own VPS; site and API on Hangar
Section titled “2026-09-09 — Search box on its own VPS; site and API on Hangar”Hangar hosts Node apps well and gives immutable releases with rollback, but provides no database and no long-running non-Node services. Meilisearch plus an embedder plus Postgres fit naturally in one Compose stack on a plain box. The apps reach it over HTTPS; the firewall admits only the Hangar server.
2026-09-09 — Ollama with bge-m3 stays the embedder
Section titled “2026-09-09 — Ollama with bge-m3 stays the embedder”The index already carries bge-m3 vectors and the code calls Ollama’s API. Alternatives were Hugging Face TEI (3–5× faster on CPU, small code change, same vectors), Meilisearch’s built-in embedder (would re-embed everything), or a paid API (different vectors, metered). Query-time embedding on a 2-vCPU box costs a few hundred milliseconds, acceptable. TEI is the upgrade if latency matters.
2026-09-09 — Search falls back to keyword; vectors can be switched off
Section titled “2026-09-09 — Search falls back to keyword; vectors can be switched off”A half-copied vector index crashed Meilisearch on every semantic query, taking keyword search down for seconds each time and rendering the site’s search page as a 404. Now: any failure on the vector path answers with keyword results marked degraded, and SEARCH_VECTORS=off forces that for all modes without a deploy. Cost: a request that silently degrades; the response says so and the UI can show it.
2026-09-09 — Cache at the edge, plus small in-process caches
Section titled “2026-09-09 — Cache at the edge, plus small in-process caches”Public pages are identical for every visitor. Cloudflare cache rule for the frontend host honouring s-maxage=600, stale-while-revalidate=3600; facets and the browse tree cached in the API for five minutes; the sitemap rebuilt hourly. Rate limit of 60 requests per 10 s per IP. Rejected: a Redis layer (a second service for what a Map does with one process).
2026-09-09 — Transcription via a queue in Payload and workers anywhere
Section titled “2026-09-09 — Transcription via a queue in Payload and workers anywhere”Options were Payload’s job queue running Python on the backend host (needs the GPU on that host), or a transcription-jobs collection claimed by workers over HTTPS with the cron secret. The second lets the GPU be a Mac at home, a Mac mini, or a rented spot instance, and survives worker preemption via heartbeats and a 30-minute stale timeout. Atomic claims use FOR UPDATE SKIP LOCKED.
2026-09-09 — Cohere ASR default model; the tashkeel fine-tune rejected
Section titled “2026-09-09 — Cohere ASR default model; the tashkeel fine-tune rejected”Measured on a 59-minute fatawa episode on Apple GPU: default model 3 min 09 s (18.8× real time), 6,447 words. The NAMAA-Space/Cohere-Speech-Tashkeel-2B fine-tune took 35 min 39 s, generated 3× the tokens, hit the token limit on 14 segments and looped on 10, and diverged from the plain transcript on 48.9 % of words. Diacritics come from a CPU pass instead (next entry).
2026-09-09 — Hamza and shadda/tanween from CAMeL Tools, no full tashkeel
Section titled “2026-09-09 — Hamza and shadda/tanween from CAMeL Tools, no full tashkeel”The owner wanted hamzas and only the necessary marks, and a light CPU path for thousands of files. CAMeL’s MLE disambiguator restores hamza, ta-marbuta and alef-maqsura at 90.3 % on 1,344 reference words, recovers 71 % of shadda and 55 % of tanween, at ~3,000 words/s on one core. Alternatives: a 350 M-parameter LLM diacritizer (slow, full vowels, unwanted) and Mishkal (weaker, no hamza).
2026-09-09 — LLM post-edit optional, guarded, off by default for bulk
Section titled “2026-09-09 — LLM post-edit optional, guarded, off by default for bulk”qwen3:8b through Ollama changed 3.8 % of words on a 20-chunk sample, mostly ta-marbuta and punctuation, with two harmful edits. Kept as an option behind a guard that rejects any line whose word content changes, at ~6 minutes per audio hour on a GPU. Recommended off for the 19,000-hour backlog and on for new content.
2026-09-09 — Meilisearch, not Elasticsearch
Section titled “2026-09-09 — Meilisearch, not Elasticsearch”Pre-dates this log; recorded here because the previous project on the search VPS ran Elasticsearch. Meilisearch runs in ~200 MB idle, has Arabic normalization built in, and supports hybrid search with user-provided vectors; Elasticsearch needs 2–4 GB just to start. The box has 3.7 GB.