返回技能列表

run-puzzle-tests

pjt222
更新于 2 days ago
3 次查看
17
2
17
在 GitHub 上查看
测试aitestingdesign

关于

This skill runs the jigsawR test suite through WSL R with three scoping options (full, filtered, or single test) and interprets the results. It's designed for post-edit verification, pre-commit checks, and debugging, ensuring renv compatibility by never using `--vanilla` mode. Developers use it to quickly identify pass/fail/skip statuses and pinpoint specific test failures.

快速安装

Claude Code

推荐
主要方式
npx skills add pjt222/agent-almanac -a claude-code
插件命令备选方式
/plugin add https://github.com/pjt222/agent-almanac
Git 克隆备选方式
git clone https://github.com/pjt222/agent-almanac.git ~/.claude/skills/run-puzzle-tests

在 Claude Code 中复制并粘贴此命令以安装该技能

技能文档

Run Puzzle Tests

Run jigsawR test suite + interpret results.

Use When

  • Post any R src edit
  • Post new puzzle type|feature
  • Pre-commit → verify nothing broke
  • Debug specific fail

In

  • Required: Scope (full|filtered|single)
  • Optional: Filter pattern ("snic", "rectangular")
  • Optional: Test file path (single mode)

Do

Step 1: Choose Scope

ScopeUse whenDuration
FullBefore commits, after major changes~2-5 min
FilteredWorking on one puzzle type~30s
SingleDebugging a specific test file~10s

→ Scope chosen by workflow: full pre-commit, filtered for type work, single for debug.

If err: unsure → default full. Slower but catches cross-type regressions.

Step 2: Create+Exec Test Script

Full:

Create /tmp/run_tests.R:

devtools::test()
R_EXE="/mnt/c/Program Files/R/R-4.5.0/bin/Rscript.exe"
cd /mnt/d/dev/p/jigsawR && "$R_EXE" -e "devtools::test()"

Filtered:

"$R_EXE" -e "devtools::test(filter = 'snic')"

Single:

"$R_EXE" -e "testthat::test_file('tests/testthat/test-snic-puzzles.R')"

→ Test out w/ pass/fail/skip.

If err:

  • NEVER --vanilla; renv needs .Rprofile
  • renv errs → renv::restore() first
  • Complex cmds fail Exit 5 → write script file

Step 3: Interpret Results

Summary line:

[ FAIL 0 | WARN 0 | SKIP 7 | PASS 2042 ]
  • PASS: Succeeded
  • FAIL: Need investigation
  • SKIP: Skipped (missing optional pkg like snic)
  • WARN: Warns (review, not blocking)

→ Summary parsed → PASS, FAIL, SKIP, WARN counts. FAIL=0 = clean.

If err: no summary → runner crashed pre-complete. Check R errs above. Output truncated → redirect: "$R_EXE" -e "devtools::test()" > test_results.txt 2>&1.

Step 4: Investigate Fails

If fail:

  1. Read msg → file, line, expected vs actual
  2. New fail or pre-existing?
  3. Assertion → read test + tested fn
  4. Error → check fn signature changed?
# Run just the failing test with verbose output
"$R_EXE" -e "testthat::test_file('tests/testthat/test-failing.R', reporter = 'summary')"

→ Root cause id'd. Real regression (fix code) or env issue (dep, path).

If err: msg unclear → add browser()|print() + re-run via testthat::test_file() for interactive debug.

Step 5: Verify Skip Reasons

Skips normal when optional deps missing:

  • snicskip_if_not_installed("snic")
  • OS-specific → skip_on_os()
  • CRAN-only → skip_on_cran()

Confirm legitimate, not masking real fails.

→ All skips accounted by legit reasons. None mask actual fails.

If err: skip suspicious → temp remove skip_if_*() + run → see if pass or hidden fail.

Check

  • All pass (FAIL=0)
  • No unexpected warns
  • Skip matches expected (only optional dep skips)
  • Test count not decreased (no accidentally removed)

Traps

  • --vanilla: Breaks renv. Never w/ jigsawR.
  • Complex -e strings: Shell escape → Exit 5. Use script files.
  • Stale pkg state: Run devtools::load_all()|document() before test if NAMESPACE changed.
  • Missing test deps: Some need Suggests pkgs. Check DESCRIPTION.
  • Parallel issues: Tests interfere → run sequential w/ testthat::test_file().

  • generate-puzzle — gen puzzles → verify behavior matches tests
  • add-puzzle-type — new types need comprehensive suites
  • write-testthat-tests — general R test patterns
  • validate-piles-notation — test PILES parse standalone

GitHub 仓库

pjt222/agent-almanac
路径: i18n/caveman-ultra/skills/run-puzzle-tests
0
agentsagentskillsai-assisted-developmentclaude-codeskillsteams

相关推荐技能

evaluating-llms-harness

测试

该Skill通过60+个学术基准测试(如MMLU、GSM8K等)评估大语言模型质量,适用于模型对比、学术研究及训练进度追踪。它支持HuggingFace、vLLM和API接口,被EleutherAI等行业领先机构广泛采用。开发者可通过简单命令行快速对模型进行多任务批量评估。

查看技能

cloudflare-cron-triggers

测试

这个Claude Skill提供了关于Cloudflare Cron Triggers的完整知识库,用于通过cron表达式定时执行Workers。它支持配置周期性任务、维护作业和自动化工作流,并能处理常见的cron触发错误。开发者可以用它来设置定时任务、测试cron处理器,并集成Workflows和Green Compute功能。

查看技能

webapp-testing

测试

该Skill为开发者提供了基于Playwright的本地Web应用测试工具集,支持自动化测试前端功能、调试UI行为、捕获屏幕截图和查看浏览器日志。它包含管理服务器生命周期的辅助脚本,可直接作为黑盒工具运行而无需阅读源码。适用于需要快速验证本地Web应用界面和交互功能的开发场景。

查看技能

finishing-a-development-branch

测试

这个Skill用于开发分支完成后的集成决策,当代码实现完成且测试通过时,它会引导开发者选择合适的工作流。它首先验证测试状态,然后提供合并、创建PR或清理等结构化选项。核心价值在于确保代码质量的同时,标准化分支收尾流程。

查看技能