Fixing search, and dropping 1.6 TB of embeddings
Almost half of transcript searches find nothing. The cause is a broken keyword search, not a missing AI feature. We can fix it with the Postgres search we already have.
Most searches come from AI agents, not people
We looked at 30 days of PostHog data. AI agents like Claude and ChatGPT (through MCP), our CLI and our own assistant make most search calls. They send short searches: 2.3 words on average.
About 6,700 people opened search in 30 days. Transcript searches are counted each time someone pauses while typing, so the real number of searches is lower. Only 5% of search sessions open the Moments (transcript) tab; most people search to open a meeting.
About 18% of agent transcript searches are in Japanese, Chinese or Korean, and half of those find nothing.
Why searches come back empty
Every transcript search goes through these steps. Three of them throw away good results.
Clean the search
Deletes every letter that isn't a–z or 0–9.
Japanese searches become empty.Look up the index
Transcripts are saved as word roots: "pricing" is saved as "price". Searches aren't.
"pricing" never matches.Find the lines
Every searched word must appear within 3 lines of each other.
Close matches get dropped.Add embeddings
A second search by meaning runs in parallel and fills some gaps.
40% of them time out at 3 s.The 3-line rule in practice
Searching for pricing discount enterprise in this meeting:
- Ali: Let's talk about the enterprise plan.
- Sara: Sure, the pricing page is ready.
- Ali: We still need to decide on timing.
- Sara: Probably next quarter.
- Ali: And the discount for yearly plans?
- Ali: Let's talk about the enterprise plan.
- Sara: Sure, the pricing page is ready.
- Ali: We still need to decide on timing.
- Sara: Probably next quarter.
- Ali: And the discount for yearly plans?
There's one more problem: only the embedding job writes the keyword index. If we simply turned embeddings off, new meetings couldn't be searched at all.
What embeddings really add
We replayed 297 real searches against production, read only. We ran each one with keyword search and with embeddings, and had an LLM judge if the results were useful.
Embeddings rescued 44 searches. In 40 of them, the searched words were in the transcript. Keyword search only missed them because of the bugs above.
A fixed keyword search does as well as embeddings
We built a test on a laptop: 232 public meetings with known answers (the QMSum dataset), plus 30 generated Japanese meetings. For each short search, we checked if the right part of the right meeting was in the top 10.
The MVP alone triples English results and lifts Japanese by half. The tie-break and agent retries add more on top.
How "rank by rare words" works
A word that shows up everywhere tells you little. A word that shows up in a few places tells you a lot. So each searched word gets a weight, based on how many lines in the account contain it.
Example: searching salesforce renewal meeting in an account with 10,000 lines.
meetingin 9,000 linesweight 0.8renewalin 120 linesweight 4.4salesforcein 15 linesweight 6.4Each 3-line window gets a score: the weights of the words it contains, divided by the weights of all searched words.
| Window contains | Score |
|---|---|
salesforce + renewal | (6.4 + 4.4) ÷ 11.6 = 93% |
renewal + meeting | (4.4 + 0.8) ÷ 11.6 = 45% |
meeting only | 0.8 ÷ 11.6 = 7% |
The first window wins even though it's missing a word. In the MVP, ties go to the newest meeting. Later, we'll break ties by how often the words appear for the window's length, which search engines call BM25. This is the same idea Google and every classic search engine started with, and it needs no AI.
Ideas we tested
| Idea | Result | Keep? |
|---|---|---|
| Rank by rare wordsAs above. Part of the MVP fix. | With the other MVP fixes: English 11% → 32%, Japanese 33% → 50%. | MVP |
| Break ties with BM25When two windows score the same, prefer the one where the words appear more often for its length. | English 32% → 36%, Japanese 50% → 55%. | Later |
| Agent retries with other wordsAn LLM wrote 3 other searches for each one, like "cost" for "pricing". We merged the results. | Japanese 55% → 70%. No change in English. | MVP |
| Word rootsSaved and searched words by their root, so "pricing" also finds "price" and "priced". | Worse in English. Unrelated words started matching. | No |
| Partial wordsMatched any word starting with the searched word: "pric" finds "price", "pricing", "prickly". | Worse in English. Too many loose matches. | No |
| Require at least half the wordsDropped any result that matched less than half of the searched words. | Broke Japanese, where one word becomes several 2-character pieces. | No |
What we're building
The core is one small function that splits text into searchable pieces. We use the same one when saving a transcript and when searching, so both sides always agree.
Japanese, Chinese and Korean don't use spaces, so we split them into overlapping pairs of characters.
- Save the index with the transcript. Not in the embedding job.
- Match all words first. If that finds too little, match most words, and skip very common ones like "the".
- Rank by rare words. Matching "Salesforce" is worth more than matching "meeting".
- Tell agents to retry. If nothing is found, try other words or another language. This is how we cover "cost" vs. "pricing" without embeddings.
The retry only helps AI agents. People using Cmd+K or mobile get the better matching, and can change their search themselves.
An index isn't always faster
Meeting search is sometimes slow: the slowest 1% take 2–3.5 seconds. We tried adding an index on a test database with 630,000 meetings, searching from an account with 500 meetings.
With the index, the database searches everyone's meetings first and throws away the ones you can't see. So we won't add it. Transcript search doesn't have this problem: the same test took 4 ms for a small account and at most 0.3 s for a large one.
Filters, dates and related meetings
These aren't needed to remove embeddings, so they come after the MVP.
A filter bug
When an agent filters transcript search by a person and a company, the person filter is silently dropped. And "acme.com" also matches "notacme.com". We reproduced both on a test database.
Dates for agents
Agents already add dates to meeting search, in 72% of MCP searches. Letting them do the same for transcripts helps a lot. Nobody picks a date by hand: the agent turns "last month" into dates.
English test set, right part in the top 10. If the dates are wrong, the agent searches again without them.
Related meetings
We want to show similar meetings on the meeting page, also without embeddings. The information we already have works better: same calendar series, the same people (rare ones count more), the same customer, the same tags. It's free, and easy to explain: "3 shared people".
Why what we already have is enough
We already run Postgres on Supabase, and it already has a word index on every transcript. The index was built wrong, but the tool is right for the job.
- Our searches are keyword searches. Agents send 2.3 words on average, usually names, products and topics. A word index is built for exactly that.
- Search stays next to the data. Who can see which meeting (sharing, workspaces, teams) is already in Postgres. A search inside the same database can't drift out of sync or leak a meeting.
- It's fast enough. 4 ms for a small account and at most 0.3 s for a large one in our test, with a table of 220,000 transcripts.
- Ranking is ours. Ranking by rare words, the idea behind Google's early search, is a few lines of our own code. No extension needed.
Why not Turbopuffer
Turbopuffer is a hosted search engine. It's fast and does both keyword and meaning search. For us it would mean:
- Copying every transcript to another company, and keeping it in sync when meetings are edited, re-transcribed or deleted.
- Rebuilding our sharing rules in its filters, and keeping them correct forever. A mistake there shows someone a meeting they shouldn't see.
- Another vendor, bill and security review.
All of that for about 1% better results on our data. We'd look again if we need real meaning-based search at large scale, or if Postgres gets too slow.
Why not LSA
LSA (latent semantic analysis) is the 1990s ancestor of embeddings. It looks at which words appear together across all meetings and turns each meeting into a list of numbers. Meetings with similar numbers are "related".
- It has the same cost as embeddings: a stored list of numbers per meeting, and a "find the closest" search.
- It has to be rebuilt as new meetings arrive, and per account, so one customer's words don't shape another's results.
- It needs words split by spaces, so it's weak in Japanese, Chinese and Korean.
- Modern embeddings beat it, and we're removing those because they only add about 1%.
For related meetings, the information we already have (same series, same people, same customer) is stronger and free.
Other options
| Option | Why not |
|---|---|
| Keep embeddings | 1.6 TB and a 16 GB service for about 1% of results. 40% of their searches time out. |
| A BM25 extension (pg_search, VectorChord) | Supabase doesn't offer one. We only need its scoring to break ties. |
| An index for meeting search | Made common words up to 9 seconds slower. |
The order of work
MVP first: the smallest version that lets us remove embeddings safely. The main line has to happen in order, since each step needs the one before it. Everything after the MVP line is improvement.
How we keep it safe
- Test before merging. Every change runs through the search test, including speed with common words.
- Fix first, remove second. Embeddings are only turned off once the new keyword search and agent retries are live.
- Watch for a week. If more searches come back empty in any language, we turn embeddings back on.
- Delete last. The table is dropped only after a backup to Cloud Storage.