SKILL·41047C

run-puzzle-tests

Name: run-puzzle-tests
Author: pjt222

pjt222

Обновлено 1 month ago

9 просмотров

Тестированиеaitestingdesign

О программе

Этот навык запускает набор тестов jigsawR через WSL R, поддерживая полные тесты, подмножества с фильтрацией по шаблонам или отдельные файлы, интерпретируя результаты прохождения/провала/пропуска. Он автоматически выявляет неудачные тесты и корректно обрабатывает зависимости renv, избегая режима `--vanilla`. Разработчикам следует использовать его после изменений в исходном коде R, при добавлении функций, перед коммитами или при отладке конкретных сбоев тестов.

Быстрая установка

Claude Code

Рекомендуется

Основной

npx skills add pjt222/agent-almanac -a claude-code

Команда плагинаАльтернативный

/plugin add https://github.com/pjt222/agent-almanac

Git клонированиеАльтернативный

git clone https://github.com/pjt222/agent-almanac.git ~/.claude/skills/run-puzzle-tests

Скопируйте и вставьте эту команду в Claude Code для установки этого навыка

Документация

Run Puzzle Tests

Run the jigsawR test suite and interpret results.

When to Use

After modifying R source code in the package
After adding a new puzzle type or feature
Before committing changes
Debugging a specific test failure

Inputs

Required: Test scope (full, filtered, or single)
Optional: Filter pattern for filtered mode (e.g. "snic", "rectangular")
Optional: Specific test file path for single mode

Procedure

Step 1: Choose Test Scope

Scope	Use when	Duration
Full	Before commits, after major changes	~2-5 min
Filtered	Working on one puzzle type	~30s
Single	Debugging a specific test file	~10s

Got: Test scope selected: full before commits, filtered for one puzzle type, single for debugging one test.

If fail: If unsure, default to full suite. Slower but catches cross-type regressions.

Step 2: Create and Execute Test Script

Full suite:

Create a script (e.g., /tmp/run_tests.R):

devtools::test()

R_EXE="/mnt/c/Program Files/R/R-4.5.0/bin/Rscript.exe"
cd /mnt/d/dev/p/jigsawR && "$R_EXE" -e "devtools::test()"

Filtered by pattern:

"$R_EXE" -e "devtools::test(filter = 'snic')"

Single file:

"$R_EXE" -e "testthat::test_file('tests/testthat/test-snic-puzzles.R')"

Got: Test output with pass/fail/skip counts.

If fail:

Do NOT use --vanilla; renv needs .Rprofile to activate
On renv errors, run renv::restore() first
For complex commands failing with Exit code 5, write to a script file

Step 3: Interpret Results

Look for the summary line:

[ FAIL 0 | WARN 0 | SKIP 7 | PASS 2042 ]

PASS: Tests succeeded
FAIL: Tests failed (need investigation)
SKIP: Tests skipped (often due to optional packages like snic)
WARN: Warnings during tests (review but not blocking)

Got: Summary line parsed for PASS, FAIL, SKIP, WARN counts. FAIL = 0 for clean run.

If fail: Without summary line, the runner crashed before completing. Check for R-level errors above. If output is truncated, redirect to file: "$R_EXE" -e "devtools::test()" > test_results.txt 2>&1.

Step 4: Investigate Failures

If tests fail:

Read the failure message — includes file, line, expected vs actual
Check if new failure or pre-existing
For assertion failures, read the test and the function tested
For error failures, check if a function signature changed

# Run failing test with verbose output
"$R_EXE" -e "testthat::test_file('tests/testthat/test-failing.R', reporter = 'summary')"

Got: Root cause of each failing test identified — regression (code fix) or environment issue (missing dep, path).

If fail: With unclear failure messages, add browser() or print() and re-run with testthat::test_file() for interactive debugging.

Step 5: Verify Skip Reasons

Skips are normal when optional dependencies are missing:

snic package tests skip with skip_if_not_installed("snic")
Tests requiring specific OS skip with skip_on_os()
CRAN-only skips with skip_on_cran()

Confirm skip reasons are legitimate, not masking real failures.

Got: All skips accounted for by legitimate reasons. No skips masking failures.

If fail: If a skip seems suspicious, temporarily remove the skip_if_*() call and run the test.

Validation

All tests pass (FAIL = 0)
No unexpected warnings
Skip count matches expected (only optional deps)
Test count has not decreased (no tests accidentally removed)

Pitfalls

Using --vanilla: Breaks renv activation. Never use it with jigsawR.
Complex -e strings: Shell escaping issues cause Exit code 5. Use script files.
Stale package state: Run devtools::load_all() or devtools::document() before testing if NAMESPACE changed.
Missing test dependencies: Some tests need suggested packages. Check DESCRIPTION Suggests field.
Parallel test issues: If tests interfere, run sequentially with testthat::test_file().

Related Skills

generate-puzzle — generate puzzles to verify behavior matches tests
add-puzzle-type — new types need comprehensive test suites
write-testthat-tests — patterns for writing R tests
validate-piles-notation — test PILES parsing independently

GitHub репозиторий

pjt222/agent-almanac

Путь: i18n/caveman-lite/skills/run-puzzle-tests

agentsagentskillsai-assisted-developmentclaude-codeskillsteams

FAQ

Frequently asked questions

What is the run-puzzle-tests skill?

run-puzzle-tests is a Claude Skill by pjt222. Skills package instructions and resources that Claude loads on demand, so Claude can perform run-puzzle-tests-related tasks without extra prompting.

How do I install run-puzzle-tests?

Use the install commands on this page: add run-puzzle-tests to Claude Code as a plugin, or clone its repository into your skills directory, then restart Claude so it picks up the skill.

What category does run-puzzle-tests belong to?

run-puzzle-tests is in the Testing category, tagged ai, testing and design.

Is run-puzzle-tests free to use?

Yes. run-puzzle-tests is listed on AIMCP and free to install. It runs inside Claude, so no separate service account is required to use the skill itself.

Похожие навыки

evaluating-llms-harness

Тестирование

Этот навык Claude запускает lm-evaluation-harness для тестирования LLM на более чем 60 стандартизированных академических задачах, таких как MMLU и GSM8K. Он предназначен для разработчиков, чтобы сравнивать качество моделей, отслеживать прогресс обучения или сообщать академические результаты. Инструмент поддерживает различные бэкенды, включая модели HuggingFace и vLLM.

Просмотреть навык

cloudflare-cron-triggers

Тестирование

Этот навык предоставляет обширные знания по реализации Cloudflare Cron Triggers для планирования запуска Workers с помощью cron-выражений. Он охватывает настройку периодических задач, заданий технического обслуживания и автоматизированных рабочих процессов, а также решение распространенных проблем, таких как неверные cron-выражения и ошибки часовых поясов. Разработчики могут использовать его для настройки планировщиков обработчиков, тестирования cron-триггеров и интеграции с Workflows и Green Compute.

Просмотреть навык

webapp-testing

Тестирование

Этот навык Claude предоставляет инструментарий на базе Playwright для тестирования локальных веб-приложений с помощью Python-скриптов. Он позволяет проводить проверку фронтенда, отладку интерфейса, создание скриншотов и просмотр логов, одновременно управляя жизненным циклом сервера. Используйте его для задач автоматизации браузера, но запускайте скрипты напрямую, вместо чтения их исходного кода, чтобы избежать загрязнения контекста.

Просмотреть навык

finishing-a-development-branch

Тестирование

Этот навык помогает разработчикам завершать готовую работу, проверяя прохождение тестов и предлагая структурированные варианты интеграции. Он направляет рабочий процесс по слиянию, созданию пул-реквестов или очистке веток после завершения реализации. Используйте его, когда ваш код готов и протестирован, чтобы систематически завершать процесс разработки.

Просмотреть навык