SKILL·DDDE8A

benchmark-and-mms-planner

Name: benchmark-and-mms-planner
Author: HeshamFS

HeshamFS

Updated 1 month ago

9 views

Testingai

About

This Claude Skill helps developers create verification and validation plans for simulation codes to ensure trustworthy results. It provides structured methodologies including manufactured solutions, benchmark problems, and refinement studies with pass/fail criteria. Use it when you need to rigorously prove a solver's correctness rather than just demonstrating plausibility.

Quick Install

Claude Code

Recommended

Primary

npx skills add HeshamFS/materials-simulation-skills -a claude-code

Plugin CommandAlternative

/plugin add https://github.com/HeshamFS/materials-simulation-skills

Git CloneAlternative

git clone https://github.com/HeshamFS/materials-simulation-skills.git ~/.claude/skills/benchmark-and-mms-planner

Copy and paste this command in Claude Code to install this skill

Documentation

Benchmark And MMS Planner

Goal

Design a verification and validation plan before trusting simulation results. The skill helps agents choose manufactured solutions, benchmark cases, refinement protocols, uncertainty checks, and pass/fail criteria.

Requirements

Python 3.10+
No external dependencies
Works on Linux, macOS, and Windows

Inputs to Gather

Input	Description	Example
PDE or model class	Governing family	`diffusion`, `elasticity`, `phase-field`
Quantity of interest	Metric to validate	`interface velocity`, `L2 temperature error`
Dimension	1, 2, or 3	`2`
Expected order	Formal discretization order	`2`
Reference availability	Analytic, benchmark, or none	`analytic`
Risk level	Cost or consequence of wrong result	`high`

Decision Guidance

Use MMS when code correctness is uncertain and an analytic solution can be injected.
Use canonical benchmarks when physical model validation matters more than code verification.
Use grid/time refinement whenever the result is used for a claim, design decision, or comparison.
Use uncertainty propagation when inputs are calibrated, noisy, or experimentally measured.

Script Outputs

scripts/benchmark_mms_planner.py emits inputs and results with:

verification_strategy
mms_plan
benchmark_cases
refinement_protocol
acceptance_criteria
warnings

Workflow

Collect the governing model, quantity of interest, and risk level.
Run benchmark_mms_planner.py --json.
Treat warnings as blockers for high-risk claims.
Convert the returned protocol into tests, simulation runs, or review checklist items.

python3 skills/verification-validation/benchmark-and-mms-planner/scripts/benchmark_mms_planner.py \
  --model diffusion \
  --quantity "L2 error in temperature" \
  --dimension 2 \
  --expected-order 2 \
  --reference analytic \
  --risk high \
  --json

Error Handling

If the dimension or expected order is invalid, stop and correct the model description.
If no reference exists, use conservation and convergence checks but do not call the result validated.

Limitations

This skill plans verification work; it does not run the solver or prove that a physical model is appropriate for an experiment.

Security

Inputs are scalar strings and finite numeric values only.
The script does not execute external solvers.
File writes are not performed.
The skill uses Bash only to run its bundled script.

References

See references/vv_patterns.md for MMS, benchmark, and uncertainty planning notes.

Version History

1.0.0: Initial benchmark and MMS planning skill.

GitHub Repository

HeshamFS/materials-simulation-skills

Path: skills/verification-validation/benchmark-and-mms-planner

agent-skillsagentscli-toolscomputational-sciencellmmaterials-science

FAQ

Frequently asked questions

What is the benchmark-and-mms-planner skill?

benchmark-and-mms-planner is a Claude Skill by HeshamFS. Skills package instructions and resources that Claude loads on demand, so Claude can perform benchmark-and-mms-planner-related tasks without extra prompting.

How do I install benchmark-and-mms-planner?

Use the install commands on this page: add benchmark-and-mms-planner to Claude Code as a plugin, or clone its repository into your skills directory, then restart Claude so it picks up the skill.

What category does benchmark-and-mms-planner belong to?

benchmark-and-mms-planner is in the Testing category, tagged ai.

Is benchmark-and-mms-planner free to use?

Yes. benchmark-and-mms-planner is listed on AIMCP and free to install. It runs inside Claude, so no separate service account is required to use the skill itself.

Related Skills

evaluating-llms-harness

Testing

This Claude Skill runs the lm-evaluation-harness to benchmark LLMs across 60+ standardized academic tasks like MMLU and GSM8K. It's designed for developers to compare model quality, track training progress, or report academic results. The tool supports various backends including HuggingFace and vLLM models.

View skill

cloudflare-cron-triggers

Testing

This skill provides comprehensive knowledge for implementing Cloudflare Cron Triggers to schedule Workers using cron expressions. It covers setting up periodic tasks, maintenance jobs, and automated workflows while handling common issues like invalid cron expressions and timezone problems. Developers can use it for configuring scheduled handlers, testing cron triggers, and integrating with Workflows and Green Compute.

View skill

webapp-testing

Testing

This Claude Skill provides a Playwright-based toolkit for testing local web applications through Python scripts. It enables frontend verification, UI debugging, screenshot capture, and log viewing while managing server lifecycles. Use it for browser automation tasks but run scripts directly rather than reading their source code to avoid context pollution.

View skill

finishing-a-development-branch

Testing

This skill helps developers complete finished work by verifying tests pass and then presenting structured integration options. It guides the workflow for merging, creating PRs, or cleaning up branches after implementation is done. Use it when your code is ready and tested to systematically finalize the development process.

View skill