About
This skill automatically repeats prompts for Haiku agents to significantly boost accuracy on structured tasks like unit testing and parsing. It works by enabling bidirectional attention during the prefill stage without adding latency. Use it automatically with Haiku for operations where precision is critical.
Quick Install
Claude Code
Recommendednpx skills add asklokesh/claudeskill-loki-mode -a claude-code/plugin add https://github.com/asklokesh/claudeskill-loki-modegit clone https://github.com/asklokesh/claudeskill-loki-mode.git ~/.claude/skills/prompt-optimizationCopy and paste this command in Claude Code to install this skill
Documentation
Prompt Optimization Skill
Overview
Automatically applies prompt repetition for Haiku agents to improve accuracy by 4-5x on structured tasks.
Research Source: "Prompt Repetition Improves Non-Reasoning LLMs" (arXiv 2512.14982v1)
When to Activate
This skill activates automatically for:
- Haiku agents executing structured tasks
- Unit test execution
- Linting and formatting
- Parsing and extraction
- List operations (find, filter, count)
How It Works
BEFORE:
prompt = "Run unit tests in tests/ directory"
AFTER (with skill):
prompt = "Run unit tests in tests/ directory\n\nRun unit tests in tests/ directory"
The repeated prompt enables bidirectional attention within the parallelizable prefill stage, improving accuracy without latency penalty.
Performance Impact
| Task Type | Without Skill | With Skill | Improvement |
|---|---|---|---|
| Unit tests | 65% accuracy | 95% accuracy | +46% |
| Linting | 72% accuracy | 98% accuracy | +36% |
| Parsing | 58% accuracy | 94% accuracy | +62% |
Latency: Zero impact (occurs in prefill, not generation)
Configuration
Enable/Disable
# Enabled by default for Haiku agents
LOKI_PROMPT_REPETITION=true
# Disable if needed
LOKI_PROMPT_REPETITION=false
Repetition Count
# 2x repetition (default)
LOKI_PROMPT_REPETITION_COUNT=2
# 3x repetition (for position-critical tasks)
LOKI_PROMPT_REPETITION_COUNT=3
Agent Instructions
When you are a Haiku agent and the task involves:
- Running tests
- Executing linters
- Parsing structured data
- Finding items in lists
- Counting or filtering
Your prompt will be automatically repeated 2x to improve accuracy. No action needed from you.
If you are an Opus or Sonnet agent, this skill does NOT apply (reasoning models see no benefit from repetition).
Metrics
Track prompt optimization impact:
.loki/metrics/prompt-optimization/
├── accuracy-improvement.json
└── cost-benefit.json
References
See references/prompt-repetition.md for full documentation.
Version: 1.0.0
GitHub Repository
Frequently asked questions
What is the prompt-optimization skill?
prompt-optimization is a Claude Skill by asklokesh. Skills package instructions and resources that Claude loads on demand, so Claude can perform prompt-optimization-related tasks without extra prompting.
How do I install prompt-optimization?
Use the install commands on this page: add prompt-optimization to Claude Code as a plugin, or clone its repository into your skills directory, then restart Claude so it picks up the skill.
What category does prompt-optimization belong to?
prompt-optimization is in the Other category.
Is prompt-optimization free to use?
Yes. prompt-optimization is listed on AIMCP and free to install.
Related Skills
This Claude Skill analyzes sports betting markets including spreads, over/unders, and prop bets by examining historical trends and situational statistics to identify value bets. It provides structured markdown output with actionable recommendations for educational purposes. Developers should use this for sports betting analysis tools while noting it's designed for entertainment/education only.
LlamaGuard is Meta's 7-8B parameter model for moderating LLM inputs and outputs across six safety categories like violence and hate speech. It offers 94-95% accuracy and can be deployed using vLLM, Hugging Face, or Amazon SageMaker. Use this skill to easily integrate content filtering and safety guardrails into your AI applications.
This Claude Skill helps developers optimize cloud costs through resource rightsizing, tagging strategies, and spending analysis. It provides a framework for reducing cloud expenses and implementing cost governance across AWS, Azure, and GCP. Use it when you need to analyze infrastructure costs, right-size resources, or meet budget constraints.
This skill quantizes LLMs to 8-bit or 4-bit precision using bitsandbytes, achieving 50-75% memory reduction with minimal accuracy loss. It's ideal for running larger models on limited GPU memory or accelerating inference, supporting formats like INT8, NF4, and FP4. The skill integrates with HuggingFace Transformers and enables QLoRA training and 8-bit optimizers.
