关于
This skill enables developers to work with Outpost Bio's Waypoint microbiome foundation models and related tools. It supports key tasks like embedding samples, fine-tuning on taxonomic data, and benchmarking models using the Compass framework. It also includes utilities for converting common microbiome data formats (like MetaPhlAn or Kraken2 outputs) for use with these models.
快速安装
Claude Code
推荐npx skills add K-Dense-AI/claude-scientific-skills -a claude-code/plugin add https://github.com/K-Dense-AI/claude-scientific-skillsgit clone https://github.com/K-Dense-AI/claude-scientific-skills.git ~/.claude/skills/waypoint-bio在 Claude Code 中复制并粘贴此命令以安装该技能
技能文档
Waypoint: Outpost Bio's Open Microbiome Foundation Models
Overview
Outpost Bio open-sourced three artefacts under Apache 2.0, described in Treloar et al., bioRxiv 2026.05.02.722381:
| Artefact | What it is | Hugging Face |
|---|---|---|
| Waypoint | GPT-2-style causal LMs over taxonomic tokens, 6M–170M params | outpost-bio/Waypoint-6m, -45m, -170m |
| Atlas | 539,308 microbiome samples scraped from MGnify (485,377 pretrain / 53,931 benchmark) | outpost-bio/Atlas |
| Compass | Eight downstream tasks over four studies | outpost-bio/Compass |
The unifying idea: a microbiome sample is a sentence. Each taxon is one token, tokens are ordered by descending abundance z-score, and the model is trained with next-token prediction. A pretrained checkpoint then supplies sample-level embeddings or a fine-tuning backbone for prediction tasks.
All of it is driven by one CLI, waypoint, with five subcommands: prepare-dataset, embed,
finetune, benchmark, pretrain.
When to use
- Embedding 16S/shotgun taxonomic profiles into fixed-size vectors for clustering, visualisation, or a downstream classifier.
- Fine-tuning a Waypoint checkpoint to predict a phenotype, treatment, or continuous readout from community composition.
- Scoring your own microbiome model against Compass so the number is comparable to the paper.
- Pretraining a taxonomic language model on Atlas or on your own corpus.
- Converting profiler output (MetaPhlAn, Kraken2/Bracken, QIIME 2, MGnify TSVs) into the input format these tools expect.
Do not reach for this when you have fewer than ~1,000 labelled samples — see Scientific caveats. A random forest on relative abundances is the better tool there, and the paper says so.
Setup
pip install waypoint-bio # installs the `waypoint` command
Atlas, Compass, and every Waypoint checkpoint are gated. Access is auto-approved, but you must click through once per repo and then authenticate:
-
Request access on each repo page you need: Waypoint-6m, Waypoint-45m, Waypoint-170m, Atlas, Compass.
-
Authenticate locally:
hf auth login # or: export HF_TOKEN=hf_...
A 401/403 from any subcommand almost always means access was never requested on that specific repo —
a token alone is not enough. Use a read-scoped token. The tokenizer loads via
trust_remote_code=True, so pin a revision if you need the remote code fixed across runs.
The waypoint data format
Everything except prepare-dataset consumes waypoint format: a .parquet / .csv / .tsv
whose rows are samples, with two aligned list-columns plus any label columns you need.
| Column | Type | Notes |
|---|---|---|
Taxa | list[str] | Full lineage strings, ;-separated: k__Bacteria; p__Firmicutes; ...; g__Lactobacillus |
Relative Abundances | list[float] | Same length as Taxa, same order |
| (any) | scalar | Targets, covariates, or a Split column |
Prefer parquet. CSV/TSV stores the lists as repr strings and round-trips through ast.literal_eval.
Give full lineages, not bare names. The tokenizer extracts the genus segment (g__) from each
lineage and falls back to the most specific higher rank when genus is missing. Bare names disable
that fallback entirely.
Workflow
1. Get your data into waypoint format
If you already have a sample × taxa (or taxa × sample) abundance matrix with lineage labels:
waypoint prepare-dataset \
--input abundance_matrix.tsv \
--metadata sample_labels.csv \
--output dataset.parquet
Orientation is auto-detected from the first column header (taxonomy, lineage, taxon, otu,
#otu id ⇒ taxa-as-rows); override with --orientation. Rows are normalised to sum to 1 unless you
pass --no_normalize, and zeros are dropped unless you pass --keep_zeros.
prepare-dataset cannot read profiler output directly — MetaPhlAn uses | separators, Kraken2
reports encode the hierarchy as indentation, and QIIME 2/SILVA prefixes the domain d__ instead of
k__ (which the tokenizer silently ignores). Use the bundled converter for those:
python scripts/profiler_to_waypoint.py \
--input merged_metaphlan.tsv --format metaphlan \
--output dataset.parquet
python scripts/profiler_to_waypoint.py \
--input reports/*.kreport --format kraken \
--output dataset.parquet
python scripts/profiler_to_waypoint.py \
--input feature-table.tsv --format qiime2 \
--output dataset.parquet
See references/data-preparation.md for every input layout, rank handling, and the d__/| gotchas.
2. Check vocabulary coverage before anything else
Waypoint's vocabulary is fixed at pretraining time from Atlas. Taxa absent from it become <unk> and
are silently dropped by waypoint embed; the paper names this as the models' main limitation. A
sample whose taxa are all out-of-vocabulary yields a degenerate [BOS][EOS] embedding.
python scripts/vocab_coverage.py --model outpost-bio/Waypoint-6m --data dataset.parquet
It reports per-sample and abundance-weighted coverage and flags samples below a threshold. Treat median abundance-weighted coverage under ~0.8 as a reason to re-examine your taxonomy labels before trusting any downstream number.
3. Embed samples
waypoint embed \
--model outpost-bio/Waypoint-6m \
--data dataset.parquet \
--output embeddings.parquet
Output is indexed by sample ID with columns dim_0 … dim_{H-1} (H = 256 for 6m, 512 for 45m,
768 for 170m). Defaults: --pooling last_token, --batch_size 32, --max_length 512, device
auto-detected (cuda → mps → cpu).
Keep --pooling last_token unless you have a reason to change it: it matches how the checkpoints
were pretrained and how benchmark and finetune pool. mean is a reasonable alternative for
unsupervised use; first_token/cls_token return the BOS position and carry little signal in a
causal LM.
4. Fine-tune on your labels
# classification
waypoint finetune \
--model outpost-bio/Waypoint-45m \
--data dataset.parquet \
--output_dir outputs/ft_disease \
--task_type classification \
--target "Disease Status" \
--config configs/finetune_classification.yaml
# regression, with a categorical covariate one-hot appended to the pooled embedding
waypoint finetune \
--model outpost-bio/Waypoint-45m \
--data dataset.parquet \
--output_dir outputs/ft_degradation \
--task_type regression \
--target "Degradation Rate" \
--covariate_column Drug \
--config configs/finetune_regression.yaml
Config paths resolve against the bundled waypoint_bio/configs/ tree, so configs/... works from
any directory without cloning.
Defaults worth overriding for small datasets: warmup_steps: 1000 (drop to ~50 so warmup finishes
before early stopping), num_epochs: 1 in the shipped configs (raise it — early stopping on
validation loss is what actually terminates training), and use_lora: true when VRAM is tight
(~1% of parameters trained; adapters are merged back before saving, so the checkpoint stays a plain
AutoModel).
Splits default to a random 80/10/10. Set split_column to a Split column whenever samples are
correlated — repeated measures, one donor sampled over time, technical replicates — or a random
split leaks and the test score is meaningless.
Outputs land in --output_dir: best_model/ (loadable by embed/benchmark),
test_metrics.json, training_log.csv + .html, and finetune_results.json.
5. Benchmark on Compass
waypoint benchmark --model outpost-bio/Waypoint-6m --output_dir outputs/benchmark
waypoint benchmark --model outputs/pretrain/best_model --tasks 1 6 --output_dir outputs/smoke
Fine-tunes a fresh head per task and writes benchmark_results.json. Classification tasks score
macro-F1; the one regression task scores R² clamped to [0, 1]; final_score is the unweighted mean
across tasks. Full task table, metric keys, and result-file schema: references/compass-benchmark.md.
6. Pretrain
waypoint pretrain \
--model_config configs/models/gpt2-45m.yaml \
--pretrain_config configs/pretraining.yaml \
--output_dir outputs/pretrain_45m
Downloads Atlas, builds a taxonomic tokenizer from the corpus, computes per-token abundance
mean/std for z-score ordering, then trains with next-token prediction and early stopping. Add
--data my_corpus.parquet to pretrain on your own waypoint-format corpus instead, and
--max_samples N for a smoke test.
Nine architectures ship, from gpt2-6m.yaml (8 layers, 256 hidden) to gpt2-170m.yaml (24 layers,
768 hidden); per-head dimension is fixed at 64 throughout. references/cli-reference.md has the
full table and every config key.
Scientific caveats
These are load-bearing. Ignoring them produces numbers that look fine and mean nothing.
- Below ~1,000 labelled examples, Waypoint underperforms a random forest on raw abundances. The paper's crossover against the RF baseline sits near 10,000 training examples. Fit the baseline first; only adopt the transformer if it wins on your data.
- Out-of-vocabulary taxa are dropped, not flagged. Every Compass dataset carries some. Run
scripts/vocab_coverage.pyand report the coverage alongside your results. - 45M, not 170M, was the best benchmark model. Pretraining loss keeps falling with scale, but downstream Compass score does not — start at 6m or 45m and only scale up if it demonstrably helps.
- Genus-level tokenisation is the default, so species-level distinctions are collapsed. Changing
taxon_rankrequires re-pretraining, not just re-tokenising. - Compositional data. Relative abundances are constrained to sum to 1; differences in one taxon induce apparent changes in others. This affects interpretation of any per-taxon attribution.
- Batch and study effects dominate microbiome data. Atlas spans MGnify pipelines v1.0–v5.0 and four sequencing modalities. Never let a study or run boundary coincide with your label boundary.
- Not a clinical or diagnostic tool. The model cards state this explicitly.
References
references/cli-reference.md— every subcommand flag, every config key, the model-size table.references/compass-benchmark.md— the eight tasks, filters, metrics,benchmark_results.jsonschema.references/data-preparation.md— waypoint format, profiler conversions, taxonomy string rules.references/python-api.md— using the tokenizer, datasets, heads, and checkpoints from Python.
Scripts
scripts/profiler_to_waypoint.py— MetaPhlAn / Kraken2 / QIIME 2 / generic lineage tables → waypoint format.scripts/vocab_coverage.py— tokenizer coverage report for a waypoint-format file.
Upstream
Code github.com/Outpost-Bio/waypoint ·
package waypoint-bio ·
paper bioRxiv 2026.05.02.722381 ·
community Waypoint Slack ·
contact [email protected].
Cite Treloar, N. J., Ur-Rehman, S., Yang, J., & Outpost Bio (2026). Learning the Language of the Microbiome with Transformers. bioRxiv. Per-artefact DOIs are listed at outpost.bio/citations.
GitHub 仓库
常见问题
什么是 waypoint-bio Skill?
waypoint-bio 是一个 Claude Skill,作者为 K-Dense-AI。Skill 将 Claude 按需加载的说明和资源打包,让 Claude 无需额外提示即可执行与 waypoint-bio 相关的任务。
如何安装 waypoint-bio?
使用本页的安装命令:将 waypoint-bio 作为插件添加到 Claude Code,或将其仓库克隆到 skills 目录,然后重启 Claude 以加载该 Skill。
waypoint-bio 属于哪个分类?
waypoint-bio 属于其他分类。
waypoint-bio 可以免费使用吗?
可以。waypoint-bio 已收录在 AIMCP,可免费安装。
相关推荐技能
LlamaGuard是Meta推出的7-8B参数内容审核模型,专门用于过滤LLM的输入和输出内容。它能检测六大安全风险类别(暴力/仇恨、性内容、武器、违禁品、自残、犯罪计划),准确率达94-95%。开发者可通过HuggingFace、vLLM或Sagemaker快速部署,并能与NeMo Guardrails集成实现自动化安全防护。
这个Claude Skill帮助开发者优化云成本,通过资源调整、标记策略和预留实例来降低AWS、Azure和GCP的开支。它适用于减少云支出、分析基础设施成本或实施成本治理策略的场景。关键功能包括提供成本可视化、资源规模调整指导和定价模型优化建议。
该Skill为开发者提供体育博彩数据分析工具,可分析盘口、大小球和特殊投注,识别价值投注机会。它整合历史数据和情景统计,生成包含时间戳的结构化Markdown报告。适用于需要快速获取博彩市场洞察的娱乐或教育类应用开发。
这个Skill使用bitsandbytes库量化大语言模型,能在GPU内存有限时通过8位或4位量化减少50-75%内存占用,同时保持精度损失最小。它支持INT8、NF4、FP4等多种量化格式,可与HuggingFace Transformers无缝集成,适用于需要部署更大模型或加速推理的场景。还提供QLoRA训练和8位优化器支持,让开发者能轻松实现高效模型压缩。
