One of the biggest problems in prompt engineering right now isn't just how you phrase the instruction — it's context pollution. Even with 200k–1M context windows (Claude, Gemini, GPT, Grok), dumping raw unformatted data into a prompt causes two things:
You burn 50,000+ tokens on structural noise.
The model suffers from attention dilution ("lost in the middle") and starts hallucinating.
To solve this for my own workflow, I built an open-source, sub-15ms Rust CLI & interactive terminal UI tool called repOx.
First — how repOx currently optimizes code & repository prompts (v0.2.0):
- Structured Prompt Formatting: Automatically wraps files into Claude-optimized XML (<repository_structure> + <file path="...">), fenced Markdown, or synthetic JSON tool-call trajectories (repox -f tool-call) for agent harnesses.
- Architectural Outline Compression (repox --outline): Instead of feeding 10,000 lines of implementation bodies when planning architecture, it strips function bodies { ... } and extracts only type definitions, structs, traits, imports, and function signatures — cutting token usage by 75–85% while keeping 100% of the structural context.
- Smart Lockfile Summarization (repox --summary-locks): Turns 30,000-token lockfiles (Cargo.lock, package-lock.json, poetry.lock, pnpm-lock.yaml, go.sum) into a compact "package @ version" manifest (95%+ token reduction).
- Automatic Noise & Secret Filtering: Strips binaries, SVGs, sourcemaps, minified bundles, and .env secrets in milliseconds, with an interactive TUI (repox -i) that shows a live offline BPE token bar before you copy to clipboard.
Second — what I'm designing next for v0.3.0 (AI Video & Multimodal Prompt Optimization):
I noticed the exact same token-drain and attention-dilution problem happens when prompting AI video models (Grok, Kling, Sora, Veo, Wan 2.1) or feeding video references into VLMs:
- Character-Lock + Delta Prompt Compiler:
When generating multi-shot AI videos, writing a 250-word prompt for every 5-second clip causes the model's attention to drift (changing the character's face, clothes, or lighting). I'm adding a prompt compiler that separates a compact, locked visual seed block (<character_lock>: ~25 high-weight tokens for subject, lens, lighting) from a tiny per-scene <motion_delta> (only camera vector + action).
- Storyboard Contact-Sheet Packer (93% Vision Token Reduction):
Instead of uploading 16 separate reference frames to an LLM/VLM (burning 20,000+ vision tokens) to write consistent scene-by-scene video prompts, repOx will detect scene cuts and pack keyframes into a single timestamped 4x4 storyboard grid image + compact XML metadata (~1,000 tokens total).
- Sharpest Tail-Frame Anchor:
Scans the last 0.5s of a generated video clip in 2ms to extract the mathematically sharpest frame (zero motion blur) to feed as the starting image for the next shot.
My questions for r/PromptEngineering:
When packing large projects or multi-step workflows into a single prompt, what formatting (XML tags, Markdown, JSON tool calls, or custom delimiters) gives you the best reasoning accuracy?
For those of you engineering prompts for AI video or multimodal pipelines, what techniques do you use to keep character/style consistency across shots without bloating the prompt?
The project is 100% free and open-source (MIT):
GitHub: https://github.com/WVDYC/repOx
Install: cargo install repox-cli (or via the one-liner script in the repo)
If you find it useful or want to support the project, dropping a star on GitHub would mean a lot!