MCP HubMCP Hub
スキル一覧に戻る

llm-inference

dave1010
更新日 Today
40 閲覧
0
GitHubで表示
デザインaidesign

について

このスキルは、OpenAI互換エンドポイントを備えたCloudflare Pages Functionsを通じてLLM推論を実現します。複数のモデルへのアクセスを提供し、gpt-oss-120bのような高性能オプションや様々なタスク向けの専門モデルを含みます。アプリケーションにLLM機能を統合する必要があり、エージェントに要件に基づいて最適なモデルを選択させたい場合にご利用ください。

クイックインストール

Claude Code

推奨
プラグインコマンド推奨
/plugin add https://github.com/dave1010/tools
Git クローン代替
git clone https://github.com/dave1010/tools.git ~/.claude/skills/llm-inference

このコマンドをClaude Codeにコピー&ペーストしてスキルをインストールします

ドキュメント

LLM Inference

The Cloudflare Pages function functions/cerebras-chat.ts provides OpenAI-compatible LLM inference. See tools/cerebras-llm-inference/index.html for a working example.

Available models

ModelMax context tokensRequests / minuteTokens / minute
gpt-oss-120b65,5363064,000
llama-3.3-70b65,5363064,000
llama3.1-8b8,1923060,000
qwen-3-235b-a22b-instruct-250765,5363064,000
qwen-3-235b-a22b-thinking-250765,5363060,000
qwen-3-32b65,5363064,000
zai-glm-4.664,00010150,000
  • llama3.1-8b is the fastest option.
  • zai-glm-4.6 is the most powerful option.
  • gpt-oss-120b remains the best all rounder.

LLMs are not just for chat: they can be used to process any string in any arbitrary way. If making a tool that requires the LLM to respond in a specific way or format then be very clear and explicit in its system prompt; eg what to include/exclude, plain/markdown formatting, length, etc.

GitHub リポジトリ

dave1010/tools
パス: .skills/llm-inference

関連スキル

content-collections

メタ

This skill provides a production-tested setup for Content Collections, a TypeScript-first tool that transforms Markdown/MDX files into type-safe data collections with Zod validation. Use it when building blogs, documentation sites, or content-heavy Vite + React applications to ensure type safety and automatic content validation. It covers everything from Vite plugin configuration and MDX compilation to deployment optimization and schema validation.

スキルを見る

creating-opencode-plugins

メタ

This skill provides the structure and API specifications for creating OpenCode plugins that hook into 25+ event types like commands, files, and LSP operations. It offers implementation patterns for JavaScript/TypeScript modules that intercept and extend the AI assistant's lifecycle. Use it when you need to build event-driven plugins for monitoring, custom handling, or extending OpenCode's capabilities.

スキルを見る

evaluating-llms-harness

テスト

This Claude Skill runs the lm-evaluation-harness to benchmark LLMs across 60+ standardized academic tasks like MMLU and GSM8K. It's designed for developers to compare model quality, track training progress, or report academic results. The tool supports various backends including HuggingFace and vLLM models.

スキルを見る

sglang

メタ

SGLang is a high-performance LLM serving framework that specializes in fast, structured generation for JSON, regex, and agentic workflows using its RadixAttention prefix caching. It delivers significantly faster inference, especially for tasks with repeated prefixes, making it ideal for complex, structured outputs and multi-turn conversations. Choose SGLang over alternatives like vLLM when you need constrained decoding or are building applications with extensive prefix sharing.

スキルを見る