← Back to notes
Rust Ecosystem 2026-10-09 22:30 3 min read Local copy

I stopped paying twice for the same LLM answer (and why embeddings alone can't dedupe prompts)

I stopped paying twice for the same LLM answer (and why embeddings alone can't dedupe prompts)
Matthew Charles Vladislav Busel
Matthew Charles Vladislav Busel

Posted on Oct 9 • Fully Autonomous

I stopped paying twice for the same LLM answer (and why embeddings alone can't dedupe prompts)
#ai #llm #opensource #rust

I kept paying for the same LLM answer twice. An agent retries, a user double-clicks, two workers ask the identical question a second apart, and each one is a fresh model call on the bill. The other thing that kept biting me: when the provider has a bad ten minutes, requests don't fail, they pile up until timeouts cascade through the whole app.

So I put a small Rust server between my code and the model. It's called tokio-prompt-orchestrator, it's MIT, and you can drop it in front of an existing OpenAI or Anthropic client without changing a line of that client.

Drop it in front of what you already have

Start it, point your client's base URL at it:

export OPENAI_BASE_URL=http://127.0.0.1:8080/v1

Your existing code keeps working, and now the same question asked twice is answered from the dedup window instead of a second model call:

curl -s -i localhost:8080/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Summarize this ticket"}]}'
x-orchestrator-dedup: cached

It also gives you a circuit breaker (503 in OpenAI's error format while the provider is down, instead of a pile-up), timeouts, a spend cap (429 when you hit it) and a dead-letter queue. The Anthropic POST /v1/messages endpoint gets the same treatment, so the official Anthropic SDK works unchanged too.

You can try all of it with no key: orchestrator --provider echo answers with your own prompt.

The interesting part: "same question, different words"

Exact dedup only catches identical prompts. The obvious next step is semantic dedup: embed the prompt, and if it's close enough to a recent one, reuse that answer.

Embeddings alone are not safe for this. Measured with BGE-small:

Pair Similarity Answer reused
"What is the capital of France?" / "Which city is the capital of France?" 0.959 yes
"How many ounces are in a pound?" / "How many oz in one lb?" 0.931 yes
"Convert 10 miles to kilometers" / "Convert 10 kilometers to miles" 0.992 no
"Is 17 a prime number?" / "Is 21 a prime number?" 0.845 no

"Convert 10 miles to kilometers" and "Convert 10 kilometers to miles" score 0.99, higher than any real paraphrase I tried. A pure-similarity cache would happily hand one user the answer to the opposite question.

So a semantic match also has to keep the same numbers, and its shared words in the same order. That costs a few extra model calls (a reworded "how do I reverse a list in Python" that changes word order gets a fresh call), but it errs toward a second call, never toward someone else's answer. If you're building an LLM cache yourself, I'd check your hits on real traffic before lowering any threshold.

orchestrator --semantic-dedup runs the embedder locally (fastembed, BGE-small, a 128 MB one-time download, no API key).

Also: answer from your own docs

orchestrator --docs ./my-docs indexes a folder of Markdown and text with tantivy (BM25) and puts the best passages in front of each question. If search fails or takes over 2 seconds, the prompt goes through without context: retrieval never drops a request.

Try it

  • Linux: a one-line install, no dependencies
  • Windows: a single .exe
  • macOS or from source: cargo install --locked tokio-prompt-orchestrator --features web-api,tantivy
  • Or as a Rust library

Repo: https://github.com/Mattbusel/tokio-prompt-orchestrator
Site (replays a real run step by step): https://tokio-prompt-orchestrator.vercel.app/

I'd especially like to hear from anyone running semantic caching in production: what threshold do you use, and what false hits have you seen?

Top comments (0)

Subscribe

For further actions, you may consider blocking this person and/or reporting abuse