SKILL·DDDE8A

benchmark-and-mms-planner

Name: benchmark-and-mms-planner
Author: HeshamFS

HeshamFS

更新日 1 month ago

9 閲覧

テストai

について

このClaudeスキルは、シミュレーションコードの検証・妥当性確認計画を作成し、信頼できる結果を保証する開発者を支援します。製造解、ベンチマーク問題、合否基準を伴う精緻化研究など、体系的な方法論を提供します。単なる妥当性の示唆ではなく、ソルバーの正確性を厳密に証明する必要がある場合にご利用ください。

クイックインストール

Claude Code

推奨

メイン

npx skills add HeshamFS/materials-simulation-skills -a claude-code

プラグインコマンド代替

/plugin add https://github.com/HeshamFS/materials-simulation-skills

Git クローン代替

git clone https://github.com/HeshamFS/materials-simulation-skills.git ~/.claude/skills/benchmark-and-mms-planner

このコマンドをClaude Codeにコピー＆ペーストしてスキルをインストールします

ドキュメント

Benchmark And MMS Planner

Goal

Design a verification and validation plan before trusting simulation results. The skill helps agents choose manufactured solutions, benchmark cases, refinement protocols, uncertainty checks, and pass/fail criteria.

Requirements

Python 3.10+
No external dependencies
Works on Linux, macOS, and Windows

Inputs to Gather

Input	Description	Example
PDE or model class	Governing family	`diffusion`, `elasticity`, `phase-field`
Quantity of interest	Metric to validate	`interface velocity`, `L2 temperature error`
Dimension	1, 2, or 3	`2`
Expected order	Formal discretization order	`2`
Reference availability	Analytic, benchmark, or none	`analytic`
Risk level	Cost or consequence of wrong result	`high`

Decision Guidance

Use MMS when code correctness is uncertain and an analytic solution can be injected.
Use canonical benchmarks when physical model validation matters more than code verification.
Use grid/time refinement whenever the result is used for a claim, design decision, or comparison.
Use uncertainty propagation when inputs are calibrated, noisy, or experimentally measured.

Script Outputs

scripts/benchmark_mms_planner.py emits inputs and results with:

verification_strategy
mms_plan
benchmark_cases
refinement_protocol
acceptance_criteria
warnings

Workflow

Collect the governing model, quantity of interest, and risk level.
Run benchmark_mms_planner.py --json.
Treat warnings as blockers for high-risk claims.
Convert the returned protocol into tests, simulation runs, or review checklist items.

python3 skills/verification-validation/benchmark-and-mms-planner/scripts/benchmark_mms_planner.py \
  --model diffusion \
  --quantity "L2 error in temperature" \
  --dimension 2 \
  --expected-order 2 \
  --reference analytic \
  --risk high \
  --json

Error Handling

If the dimension or expected order is invalid, stop and correct the model description.
If no reference exists, use conservation and convergence checks but do not call the result validated.

Limitations

This skill plans verification work; it does not run the solver or prove that a physical model is appropriate for an experiment.

Security

Inputs are scalar strings and finite numeric values only.
The script does not execute external solvers.
File writes are not performed.
The skill uses Bash only to run its bundled script.

References

See references/vv_patterns.md for MMS, benchmark, and uncertainty planning notes.

Version History

1.0.0: Initial benchmark and MMS planning skill.

GitHub リポジトリ

HeshamFS/materials-simulation-skills

パス: skills/verification-validation/benchmark-and-mms-planner

agent-skillsagentscli-toolscomputational-sciencellmmaterials-science

FAQ

Frequently asked questions

What is the benchmark-and-mms-planner skill?

benchmark-and-mms-planner is a Claude Skill by HeshamFS. Skills package instructions and resources that Claude loads on demand, so Claude can perform benchmark-and-mms-planner-related tasks without extra prompting.

How do I install benchmark-and-mms-planner?

Use the install commands on this page: add benchmark-and-mms-planner to Claude Code as a plugin, or clone its repository into your skills directory, then restart Claude so it picks up the skill.

What category does benchmark-and-mms-planner belong to?

benchmark-and-mms-planner is in the Testing category, tagged ai.

Is benchmark-and-mms-planner free to use?

Yes. benchmark-and-mms-planner is listed on AIMCP and free to install. It runs inside Claude, so no separate service account is required to use the skill itself.