SKILL·581F07

speech-to-text

Name: speech-to-text
Author: NoizAI

NoizAI

更新日 1 month ago

11 閲覧

518

GitHubで表示

メタword

について

このスキルは、'transcribe'（文字起こし）や 'speech to text'（音声テキスト変換）などのキーワードで起動し、オーディオ/ビデオファイルをテキストに書き起こします。多言語対応の文字起こし、話者識別、キャプション用のタイムスタンプ生成をサポートしています。開発者は、自動検出機能を備えたこのスキルを利用して、メディアファイルから音声コンテンツを抽出することができます。

クイックインストール

Claude Code

推奨

メイン

npx skills add NoizAI/skills -a claude-code

プラグインコマンド代替

/plugin add https://github.com/NoizAI/skills

Git クローン代替

git clone https://github.com/NoizAI/skills.git ~/.claude/skills/speech-to-text

このコマンドをClaude Codeにコピー＆ペーストしてスキルをインストールします

ドキュメント

speech-to-text

Transcribe any audio file to text. Supports multilingual auto-detection, timestamps, and speaker labels.

Triggers

transcribe / transcript / transcription
speech to text / STT / audio to text
what does this audio say / convert audio
转录 / 语音转文字 / 识别音频

Quick Start

# Transcribe with auto language detection
python3 skills/speech-to-text/scripts/stt.py audio.mp3

# Specify language explicitly
python3 skills/speech-to-text/scripts/stt.py interview.wav --language en

# Save transcript to file
python3 skills/speech-to-text/scripts/stt.py podcast.m4a -o transcript.txt

# Output full JSON (with timestamps and speaker labels)
python3 skills/speech-to-text/scripts/stt.py meeting.wav --json -o result.json

Arguments

Argument	Default	Description
`file`	required	Audio file to transcribe (mp3, wav, m4a, ogg, flac, aac, webm). Max 50 MB, max 10 min.
`--language` / `-l`	auto-detect	BCP-47 language code (e.g. `en`, `zh`, `ja`). Omit to auto-detect.
`--output` / `-o`	stdout	Path to save transcript text (or JSON if `--json` is set).
`--json`	off	Output full JSON response with timestamps and speaker labels.
`--api-key`	from env/config	Noiz API key (overrides stored key).

Output Format

Without --json, only the transcript text is printed:

Hello, welcome to today's podcast. We have a special guest joining us...

With --json, the full structured response is printed:

{
  "language": "en",
  "transcript": "Hello, welcome to today's podcast...",
  "duration": 42.5,
  "segments": [
    {"text": "Hello, welcome to today's podcast.", "start": 0.0, "end": 3.2, "spk": 0},
    {"text": "We have a special guest joining us.", "start": 3.5, "end": 6.1, "spk": 0}
  ]
}

Supported Languages

Common codes: en (English), zh (Chinese), ja (Japanese), ko (Korean), es (Spanish), fr (French), de (German), pt (Portuguese), ru (Russian), ar (Arabic). Omit --language to auto-detect.

Configuration

# Save your API key once
python3 skills/speech-to-text/scripts/stt.py config --set-api-key YOUR_KEY

# Or set via environment variable
export NOIZ_API_KEY=YOUR_KEY

Get your API key at developers.noiz.ai.

Pricing

Billed at $0.0006 per second of audio. A 10-minute file costs ~$0.36. New accounts include 10,000 free TTS characters; STT is billed separately.

Security & data disclosure

Credential storage: API key is saved to ~/.config/noiz/api_key (permissions 0600). NOIZ_API_KEY env var is also supported.
Network calls: The audio file is uploaded to https://noiz.ai/v1/speech-to-text for transcription. No data is sent until you run the command.
File limits: Max 50 MB per file, max 10 minutes (600 seconds) of audio.

Requirements

requests package: pip install requests
Get your API key at developers.noiz.ai

GitHub リポジトリ

NoizAI/skills

パス: skills/speech-to-text

FAQ

Frequently asked questions

What is the speech-to-text skill?

speech-to-text is a Claude Skill by NoizAI. Skills package instructions and resources that Claude loads on demand, so Claude can perform speech-to-text-related tasks without extra prompting.

How do I install speech-to-text?

Use the install commands on this page: add speech-to-text to Claude Code as a plugin, or clone its repository into your skills directory, then restart Claude so it picks up the skill.

What category does speech-to-text belong to?

speech-to-text is in the Meta category, tagged word.

Is speech-to-text free to use?

Yes. speech-to-text is listed on AIMCP and free to install. It runs inside Claude, so no separate service account is required to use the skill itself.

speech-to-text

について

クイックインストール

Claude Code

ドキュメント

speech-to-text

Triggers

Quick Start

Arguments

Output Format

Supported Languages

Configuration

Pricing

Security & data disclosure

Requirements

GitHub リポジトリ

Frequently asked questions

What is the speech-to-text skill?

How do I install speech-to-text?

What category does speech-to-text belong to?

Is speech-to-text free to use?

関連スキル