Compress a prompt
Remove filler words from long prompts so they cost less. Nothing is rewritten, and everything runs on this Mac.
Smaller by
—
Tokens
counted with cl100k_base
Saved per call
—
—
Saved per month
—
—
Time
—
—
Paste a long prompt, document, e-mail or transcript
Works best on long context (500+ tokens). You can also drop a .txt or .md file here.
or try a sample
How it works
- 1Add your textPaste it, open a file or pick a sample.
- 2Choose a strengthBalanced (keep 50%) is a good start.
- 3Compress and copyThe shorter text appears here, ready to use.
Words are removed, never rewritten. See an example →
Tips for this result
- Characters
- —
- Words
- —
- Tokens kept
- —
- Kept exactly
- —
| Original | Compressed | You save | |
|---|---|---|---|
| Per call | — | — | — |
| Per 1,000 calls | — | — | — |
| Per month | — | — | — |
| Per year | — | — | — |
Input tokens only; output tokens don't change. Prices come from pricing.json (placeholders until you edit them).
Ask a local AI (Ollama) the same question about the original and the compressed text, then compare the answers. Nothing leaves this machine.
Not checked yetAnswer similarity
—
Shared words between the two answers (ROUGE-1). A rough guide: always read both answers.
From the compressed text
Compare rates
Compress your text at several strengths in one go and see how much each saves. Pick the smallest result that still reads well.
Add text on the Compress page, then run the comparison.
| Keep | Tokens | Smaller | Ratio | Time | Actions |
|---|
Preview
first 600 characters · click a row or point to switchBatch files
Compress many .txt or .md files at once with your current settings. Files are read in this browser and processed on this machine.
| File | Tokens | After | Smaller | Time | Status | Actions |
|---|
History
Your last 10 compressions, saved only in this browser. Click one to open it again.
No compressions yet. They appear here automatically.
Guide
How PromptShrink works, when to use it, and how to read the numbers.
How it works
PromptShrink uses LLMLingua-2, a small model from Microsoft. It reads your text, gives every word a score for how much it matters, and removes the lowest-scoring words. It never rewrites or paraphrases: what's left is your own text, word for word, minus the filler.
Follow one sentence through the app. Press Play or click a step.
-
Browser
1
Paste text & settings
Keep 50% · keep numbers exactly
Okay, so I think the total budget for 2026 is $4.2M, right? -
Server
2
Set aside protected parts
Numbers, links, code, tags
Okay, so I think the total budget for 2026 is $4.2M, right? -
Server · GPU
3
Score every token
LLMLingua-2 on Apple GPU
-
Server
4
Keep the top 50%
Low scores are dropped
Okay, so Ithinkthetotal budgetfor2026is$4.2M, right? -
Server
5
Rejoin & tidy
Protected parts back in place
Okay think total budget 2026 $4.2M? -
Server
6
Count & price
tiktoken · pricing.json
21 → 13 tokens
38% smaller -
Browser
7
Results
Text, removed words, savings, tips
Okay think total budget 2026 $4.2M?
Press Play to walk through the seven steps.
Before
Maria: Okay, so I think everybody is here now, so let's go ahead and get started. Thanks everyone for joining today.
After (keep 50%)
Maria here now started. Thanks everyone for joining today.
The model was trained by asking GPT-4 to compress thousands of meeting transcripts, then learning to copy its choices. The result often reads like a telegram, but large language models understand it almost as well as the original.
Try it live
Uses the real Small modelEdit the text or drag the slider. Struck-through words are the ones the model removes; highlighted parts are kept exactly.
When to use it
Works well for
- Long background context: documents, reports, manuals
- Meeting transcripts and chat logs
- E-mails and support tickets
- Retrieved chunks in RAG pipelines
- Anything sent many times, where savings add up
Be careful with
- Short instructions: little to gain, easy to break
- Code, IDs, numbers and exact names (use “Keep exactly”)
- Legal or medical text where every word matters
- Very aggressive settings (keeping less than 30%)
Settings explained
- Strength (keep %)
- The share of tokens to keep, from 10% to 90%. 50% makes the text about 2× smaller. Levels:
- Token budget
- Give a maximum token count instead, e.g. to fit a context window. PromptShrink works out the rate for you. If the text already fits, nothing is removed.
- Model
- Small (BERT-base, ~710 MB) is fast and good for most text. Large (XLM-RoBERTa-large, ~2.2 GB) makes slightly better choices and is slower. Both handle many languages.
- Keep exactly
- Finds numbers, links and code automatically and keeps them exactly as written.
- Live update
- After the first compression, changing a setting re-runs it straight away, so you can drag the strength slider and watch the result change.
- Tidy spacing
- LLMLingua rebuilds text with a space between every piece (“( Head of Product )”). Tidy spacing fixes that, which also saves a few tokens. On Settings.
- Protected tokens
- Single characters or words that are never removed. The defaults keep line breaks and sentence punctuation so the structure survives. On Settings.
- Settings file
- Port, Ollama model, defaults and limits live in
.envin the app folder. Edit it, then run./promptshrink.sh restart. The values in effect are listed on Settings.
Protecting text
Select text in the Original box and click Protect. It's wrapped in LLMLingua's tag and kept exactly:
<llmlingua, compress=False>Answer in JSON with keys id and total.</llmlingua>
You can also give one section its own rate, for example compress background more than the main text:
<llmlingua, rate=0.3>Long background section…</llmlingua>
Tags are removed from the output and from token counts. Sections can't be nested.
Reading the numbers
- Token
- The unit LLMs read and bill by, about ¾ of an English word. Counts here use OpenAI's
cl100k_basetokenizer so runs are comparable; your provider may count slightly differently. - Smaller by
- How much smaller the text got: 50% means half the tokens are gone.
- Ratio
- Original ÷ compressed. 2.0x means half the size.
- Tokens kept
- The share actually kept. It can differ from your setting: protected parts are never compressed, and the model rounds to whole words.
- Cost
- Input tokens × price per million tokens from
pricing.json. Output tokens aren't affected by prompt compression.
Keyboard shortcuts
Use it from code
The app has a local HTTP API. This request uses your current settings:
Full interactive API docs: /api/docs
Privacy
- Compression runs entirely on this machine. Your text is never sent to an external service.
- The quality check talks only to Ollama on your own machine. Ollama cloud models are hidden and refused.
- Internet is used only to download the models once (from Hugging Face) and to load the chart library for Compare rates. Neither request contains your text.
- History and settings are stored in this browser's local storage. Clear them any time on Settings.
Settings
Preferences are saved in this browser. Server settings come from .env.
General
Appearance
Follow the system or pick one.
Compression details
Comma-separated, never removed. \n = new line, \, = comma.
Models & system
Loading…
Server settings in effect
From .env. Edit the file, then run ./promptshrink.sh restart.
Reset
Settings
Back to the defaults from .env (strength, model, protection, prices).
History
Delete the saved compressions from this browser.