关于
This skill helps developers plan and run moderated or unmoderated usability tests on prototypes or existing designs to find issues before launch. It handles test design, task scripts, moderation, and findings synthesis. Use it to validate designs, improve task completion, and ensure real users can successfully use what you've built.
快速安装
Claude Code
推荐npx skills add rampstackco/claude-skills -a claude-code/plugin add https://github.com/rampstackco/claude-skillsgit clone https://github.com/rampstackco/claude-skills.git ~/.claude/skills/usability-testing在 Claude Code 中复制并粘贴此命令以安装该技能
技能文档
Usability Testing
Plan and run tests that find usability problems before users hit them in production. Stack-agnostic. Tool-agnostic.
This skill is for testing existing designs or prototypes. For broader discovery research, use ux-research. For conversion testing in production, use cro-optimization.
When to use
- Before launching a new flow or major redesign
- After a redesign to verify it doesn't introduce new problems
- When analytics show drop-off but you don't know why
- When customer support tickets pattern around specific UI areas
- Pre-launch user validation
- Comparing two design directions
When NOT to use
- Discovery / generative research (use
ux-research) - Live conversion optimization (use
cro-optimization) - Mapping the broader experience (use
journey-mapping) - Pure quantitative measurement (use
analytics-strategy)
Required inputs
- The design or prototype to test (functional or near-functional)
- Specific tasks users would do
- The audience (who should be tested)
- Testing infrastructure (moderated tool, unmoderated tool, in-person setup)
The framework: 5 phases
1. Define what to test
Don't test the whole product. Test specific tasks.
Task selection criteria:
- The task represents a real user goal (not "click around and explore")
- The task has a clear start and end
- The task is achievable in 2 to 10 minutes
- The task is one of: most common, most strategic, most problematic
Examples of testable tasks:
"You want to find a contractor near you who can install a fence. Show me how you'd do that on this site."
"You're a first-time visitor. You want to understand if this product fits your needs. Walk me through how you'd evaluate it."
"Your team needs a new tool to manage projects. Use this site to figure out which plan is right for a 12-person team."
Task framing rules:
- State the user goal, not the system action ("find a place to stay" not "click the search button")
- Provide context (why are you doing this?)
- Don't reveal the path
- Don't use product terminology in the task framing
2. Choose moderated or unmoderated
Moderated (live, with researcher):
- Researcher observes and probes in real time
- Best for early-stage prototypes, complex tasks, novel concepts
- Higher cost, smaller sample (5 to 8 participants typical)
- Catches surprises and probe deeper
Unmoderated (recorded, asynchronous):
- Participant completes alone, often via tool (UserTesting, Maze, Lookback)
- Best for stable designs, simple tasks, larger sample
- Lower cost, larger sample (15 to 30 participants typical)
- Catches patterns at scale, less depth per session
For most teams: moderated for early/critical decisions, unmoderated for ongoing validation.
3. Recruit
Target audience - not just convenience.
Recruit criteria:
- Match real users (target audience, not just "anyone")
- Mix of experience levels with the product (new and existing if applicable)
- Mix of relevant device types (mobile, desktop, tablet if relevant)
- Exclude friends, family, employees
Sample size:
- Moderated: 5 to 8 participants (Nielsen's "5 users find 85% of usability issues" for the most common segment)
- Unmoderated: 15 to 30 participants (more participants compensate for less probing)
- Multi-segment testing: 5 to 8 per segment
4. Run the test
Pre-task setup:
- Confirm recording works
- Brief participant (purpose, anonymity, recording, "no wrong answers")
- Get verbal consent
- Have participant share screen if remote
Moderated session structure:
- Warm-up (2 to 3 min). Easy questions to put participant at ease.
- Pre-test questions (3 to 5 min). Background context, current behavior with similar products.
- Task 1 (5 to 10 min). Describe task. Have participant attempt while thinking aloud.
- Post-task questions (1 to 2 min). What was easy/hard? Anything confusing?
- Repeat for tasks 2, 3, 4 (typically 3 to 5 tasks per 60-minute session).
- Overall debrief (5 to 10 min). General reactions, comparisons to alternatives, anything else.
- Close (2 min).
Moderation principles:
- Encourage think-aloud ("What's going through your mind?")
- Don't help unless they're truly stuck (and even then, only after a long pause)
- Don't lead ("Are you looking for the menu?" - bad)
- Note where they hesitate, scroll, or backtrack
- Note their language vs the product's language
- Note emotional reactions
Anti-patterns:
- Talking too much (researcher should talk maybe 20% of the time)
- Defending the design when participants struggle
- Helping prematurely
- Asking participants to predict their future behavior
- Treating participant suggestions as features ("Users want X" - test demand for X separately)
5. Synthesize and report
Patterns across participants are signal. Single-participant complaints are weaker (but worth investigating).
Synthesis steps:
- Issue inventory. Every issue observed, with which participant, which task, severity.
- Cluster. Issues that are the same root problem.
- Severity.
- Critical: Blocks task completion. Most users hit this.
- Major: Significantly slows task. Many users hit this.
- Minor: Friction. Some users hit this. Workaround exists.
- Cosmetic: Polish. Doesn't affect task.
- Recommendations. For each issue, propose specific fixes.
- Prioritize. By severity and effort.
Report structure:
# Usability Test: [Design / flow]
## Summary
[2 to 3 paragraphs covering: what was tested, headline findings, top 3 priorities]
## Method
[Moderated/unmoderated, sample size, audience, dates, tasks]
## Critical findings
[Each with description, frequency, supporting evidence (quotes/clips), recommendation. Where evidence was not obtained, state the gap per the data-availability rule]
## Major findings
[Same structure]
## Minor findings
[Brief]
## Cosmetic findings
[Briefest]
## What worked well
[Calibration: capture successes too]
## Recommendations
[Prioritized list with effort estimates]
## Next steps
[Test re-run schedule, design iteration plan]
Workflow
- Define the goals. What decisions hinge on this? What tasks matter most?
- Design tasks. 3 to 5 specific, realistic, goal-framed tasks.
- Choose moderated vs unmoderated. Match to stage and depth needed.
- Recruit. Specific to audience.
- Pilot. 1 to 2 sessions before main batch. Refine tasks if needed.
- Run. Follow the protocol. Stay disciplined.
- Synthesize during, not just after. Patterns emerge by session 4 or 5.
- Report. Multiple formats - written report + highlight clips.
- Track fixes. Every critical issue should have an owner and date.
- Re-test after fixes. Verify the fix worked, didn't introduce new issues.
Failure patterns
- Testing the whole product instead of specific tasks. Vague results.
- Tasks that reveal the path. ("Click the menu and find...")
- Friends and family as participants. Biased, not representative.
- Researcher leading the participant. Findings reflect the researcher.
- Defending the design when participants struggle. Misses real issues.
- Helping too quickly. Participant doesn't experience the friction.
- Treating participant suggestions as features. Users solve their problem; product team designs the solution.
- One participant = data point. A single strong opinion isn't a finding.
- Skipping severity scoring. All findings treated equally; team can't prioritize.
- Reports no one reads. Highlight clips and live walkthroughs work better than 80-page decks.
- Testing once, never re-testing. Fixes that introduce new problems go undetected.
Output format
Default outputs:
- Test plan (before testing) -
usability-test-plan-[topic].md - Task script (per session) -
usability-tasks-[topic].md - Findings report (after synthesis) -
usability-findings-[topic].md - Highlight clips (separately produced)
If required data is unavailable
This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.
Reference files
references/task-script-patterns.md- Task framing patterns by common product type, with good and bad examples.
GitHub 仓库
常见问题
什么是 usability-testing Skill?
usability-testing 是一个 Claude Skill,作者为 rampstackco。Skill 将 Claude 按需加载的说明和资源打包,让 Claude 无需额外提示即可执行与 usability-testing 相关的任务。
如何安装 usability-testing?
使用本页的安装命令:将 usability-testing 作为插件添加到 Claude Code,或将其仓库克隆到 skills 目录,然后重启 Claude 以加载该 Skill。
usability-testing 属于哪个分类?
usability-testing 属于元分类。
usability-testing 可以免费使用吗?
可以。usability-testing 已收录在 AIMCP,可免费安装。
相关推荐技能
Content Collections 是一个 TypeScript 优先的构建工具,可将本地 Markdown/MDX 文件转换为类型安全的数据集合。它专为构建博客、文档站和内容密集型 Vite+React 应用而设计,提供基于 Zod 的自动模式验证。该工具涵盖从 Vite 插件配置、MDX 编译到生产环境部署的完整工作流。
这个Claude Skill为开发者提供完整的Polymarket预测市场开发支持,涵盖API调用、交易执行和市场数据分析。关键特性包括实时WebSocket数据流,可监控实时交易、订单和市场动态。开发者可用它构建预测市场应用、实施交易策略并集成实时市场预测功能。
该Skill帮助开发者创建OpenCode插件,用于接入命令、文件、LSP等25+种事件。它提供了插件结构、事件API规范和JavaScript/TypeScript实现模式,适合需要拦截操作、扩展功能或自定义事件处理的场景。开发者可通过它快速构建响应式模块来增强OpenCode AI助手的能力。
SGLang是一个专为LLM设计的高性能推理框架,特别适用于需要结构化输出的场景。它通过RadixAttention前缀缓存技术,在处理JSON、正则表达式、工具调用等具有重复前缀的复杂工作流时,能实现极速生成。如果你正在构建智能体或多轮对话系统,并追求远超vLLM的推理性能,SGLang是理想选择。
