SKILL·0F29EF

text-to-music

NoizAI
Updated 10 days ago
6 views
526
78
526
View on GitHub
Metageneral

About

This skill generates complete songs with vocals from text descriptions or lyrics, and can create cover versions in new styles. It handles requests for original compositions, jingles, and style transfers, but not instrumental tracks or sound effects. Use it when users want to turn lyrics into music, remake existing songs, or produce vocal-based audio from prompts.

Quick Install

Claude Code

Recommended
Primary
npx skills add NoizAI/skills -a claude-code
Plugin CommandAlternative
/plugin add https://github.com/NoizAI/skills
Git CloneAlternative
git clone https://github.com/NoizAI/skills.git ~/.claude/skills/text-to-music

Copy and paste this command in Claude Code to install this skill

Documentation

text-to-music

Two ways to make a song with vocals + accompaniment, saved as MP3. Every request produces two variants.

  • Generate — a new song from a music description plus lyrics.
  • Cover — re-sing an existing song in a new style, keeping its melody. Lyrics can be the original (auto-recognized) or new ones.

Triggers

  • make a song / generate music / text to music / compose
  • turn these lyrics into a song / sing this / jingle / theme song / birthday song
  • cover this song / make a jazz version / remake in another style
  • 写歌 / 生成音乐 / 作曲 / 唱首歌 / 歌词谱曲 / 翻唱 / 改编 / 换个风格唱

Generate — new song from text

A request needs a prompt (the music description) and exactly one lyrics source:

# Let the server write lyrics from a theme
python3 skills/text-to-music/scripts/t2m.py "upbeat indie pop, bright guitars, 120 BPM" \
  --lyrics-prompt "a song about summer road trips with friends"

# Use your own lyrics (inline or from a file)
python3 skills/text-to-music/scripts/t2m.py "mandopop ballad, piano and strings" \
  --lyrics $'[Verse]\n窗外的雨停了\n[Chorus]\n我还在等你' --vocal-gender Female --title "雨停了"
python3 skills/text-to-music/scripts/t2m.py "lo-fi hip hop, mellow, vinyl crackle" \
  --lyrics-file song.txt --duration 120 -o chill.mp3

# Model the vocal/style on a reference song (local file ≤20 MB or public URL)
python3 skills/text-to-music/scripts/t2m.py "dreamy synth-pop" --lyrics-file song.txt --ref-audio ./ref.mp3

Output: song_v1.mp3 and song_v2.mp3 (or <stem>_v1.mp3 / <stem>_v2.mp3 with -o). When lyrics were written by the server, they are also saved to <stem>_lyrics.txt.

ArgumentDefaultDescription
promptrequiredMusic description: genre, mood, instruments, tempo (≤1500 chars)
--lyrics / -l—Lyrics text; [Verse], [Chorus], [Bridge] markers help structure
--lyrics-file—UTF-8 .txt file with lyrics (≤64 KB)
--lyrics-prompt—Theme for server-side lyric writing
--tags—Style tags, comma-separated (≤50 chars)
--negative-tags—Styles to avoid
--duration / -dautoTarget length hint: 60, 120, or 180 seconds. A hint, not a guarantee.
--ref-audio—Reference song, local path or public URL

--lyrics, --lyrics-file, and --lyrics-prompt are mutually exclusive; one is required.

Cover — re-sing an existing song

Needs the source song, lyrics, and a description of the new sound (--style and/or --prompt):

# Same lyrics, new style: recognize lyrics from the original, then cover it
python3 skills/text-to-music/scripts/t2m.py cover original.mp3 \
  --style "jazz, smoky female vocal" --recognize-lyrics -o jazz_cover.mp3

# New lyrics over the original melody, looser arrangement
python3 skills/text-to-music/scripts/t2m.py cover https://example.com/song.mp3 \
  --prompt "acoustic folk, fingerpicked guitar, intimate" --lyrics-file new_lyrics.txt \
  --adherence main_melody --vocal-gender Male

# Recognize lyrics only (to review or edit before covering)
python3 skills/text-to-music/scripts/t2m.py lyrics original.mp3 -o lyrics.txt

Recommended flow when the user wants to tweak lyrics: run lyrics → edit the file → cover --lyrics-file.

ArgumentDefaultDescription
sourcerequiredOriginal song: local file (≤100 MB; mp3/wav/m4a/aac/flac/ogg) or public URL
--style / -s—Short style tags, e.g. "rock, male vocal" (≤50 chars)
--prompt / -p—Longer description of the cover's sound (≤1500 chars)
--adherencehighhigh = follow the original arrangement closely; main_melody = keep only the main melody, more freedom
--lyrics / -l—Lyrics to sing
--lyrics-file—UTF-8 .txt file with lyrics (≤64 KB)
--recognize-lyricsoffTranscribe the source's lyrics first and use them
--output / -ocover.mp3Output path; variants get _v1 / _v2 suffixes

At least one of --style / --prompt is required; exactly one lyrics source is required.

Shared Arguments

ArgumentDefaultDescription
--title—Song title (≤80 chars)
--vocal-gender—Male or Female
--vocal-timbre—Free-form voice description, e.g. "husky warm alto" (≤50 chars)
--output / -osong.mp3 (cover.mp3 for covers)Output path; variants get _v1 / _v2 suffixes
--no-waitoffSubmit, print the job ID(s), then exit
--poll-interval10Seconds between status checks
--timeout1800Max seconds to wait (the job keeps running server-side)
--api-keyfrom env/configNoiz API key (overrides stored key)

Resuming / Checking Jobs

Generation prints two gen_product_ids; a cover prints one task_id. If the wait times out or you used --no-wait, resume with:

python3 skills/text-to-music/scripts/t2m.py status <id_v1> <id_v2> -o song.mp3
python3 skills/text-to-music/scripts/t2m.py status --cover <task_id> -o cover.mp3
python3 skills/text-to-music/scripts/t2m.py status --cover <task_id> --no-wait   # just print status

Status goes submitted → running → succeeded / failed. Covers also report a stage (analysis runs before singing, so covers take longer). Download URLs expire after ~24h; re-run status for fresh ones.

Writing Good Prompts

  • Prompt / style = the sound: genre + mood + instrumentation + tempo. "cinematic orchestral pop, soaring strings, 90 BPM, hopeful".
  • Lyrics: keep sections short and label them. Chinese, English, and mixed lyrics all work.
  • Short songs: for jingles or greetings, pass -d 60 and write only a verse + chorus.
  • Covers: pick high to keep the cover recognizable; pick main_melody for a bigger genre change.
  • Not satisfied? Run again — each run gives two fresh variants.

Cost

  • Songs and covers are billed once per job (not per variant): ceil(longest variant seconds) × 15 credits, charged after both variants succeed. A 2-minute song costs about 1,800 credits. If subscription credits don't cover the whole job, it is billed as pay-as-you-go at $0.00015/s. The script prints the charged amount on completion.
  • Lyrics recognition (lyrics / --recognize-lyrics): 100 credits (or $0.001 pay-as-you-go), charged only when vocals are detected.

Rate limits: 5 submissions per minute per key per endpoint; status polling 60 per minute.

Limitations

  • Pure instrumental music is not supported — lyrics are always required.
  • --duration is only a hint; actual length may differ. Covers follow the source's length.
  • Prompts or lyrics containing sensitive content are rejected (code=400).

Configuration

python3 skills/text-to-music/scripts/t2m.py config --set-api-key YOUR_KEY
# or
export NOIZ_API_KEY=YOUR_KEY

Get your API key at developers.noiz.ai.

Requirements

  • Python 3.6+
  • requests package: uv pip install requests
  • Noiz API key from developers.noiz.ai

Security & Data Disclosure

  • API key: Stored in ~/.config/noiz/api_key (permissions 0600) or via NOIZ_API_KEY env variable.
  • Network: Prompts, lyrics, and any reference/source audio (file upload or URL) are sent to https://noiz.ai/v1/text-to-music and https://noiz.ai/v1/text-to-music/cover*; job status is polled from the same endpoints. No other data is transmitted.
  • Output: Generated MP3s (and lyrics text files) are downloaded from signed Noiz storage URLs and saved to the output path. No other files are modified.

GitHub Repository

NoizAI/skills
Path: skills/text-to-music
0
FAQ

Frequently asked questions

What is the text-to-music skill?

text-to-music is a Claude Skill by NoizAI. Skills package instructions and resources that Claude loads on demand, so Claude can perform text-to-music-related tasks without extra prompting.

How do I install text-to-music?

Use the install commands on this page: add text-to-music to Claude Code as a plugin, or clone its repository into your skills directory, then restart Claude so it picks up the skill.

What category does text-to-music belong to?

text-to-music is in the Meta category.

Is text-to-music free to use?

Yes. text-to-music is listed on AIMCP and free to install.

Related Skills

content-collections
Meta

This skill provides a production-tested setup for Content Collections, a TypeScript-first tool that transforms Markdown/MDX files into type-safe data collections with Zod validation. Use it when building blogs, documentation sites, or content-heavy Vite + React applications to ensure type safety and automatic content validation. It covers everything from Vite plugin configuration and MDX compilation to deployment optimization and schema validation.

View skill
polymarket
Meta

This skill enables developers to build applications with the Polymarket prediction markets platform, including API integration for trading and market data. It also provides real-time data streaming via WebSocket to monitor live trades and market activity. Use it for implementing trading strategies or creating tools that process live market updates.

View skill
creating-opencode-plugins
Meta

This skill helps developers create OpenCode plugins that hook into 25+ event types like commands, files, and LSP operations. It provides the plugin structure, event API specifications, and implementation patterns for JavaScript/TypeScript modules. Use it when you need to intercept, monitor, or extend the OpenCode AI assistant's lifecycle with custom event-driven logic.

View skill
sglang
Meta

SGLang is a high-performance LLM serving framework that specializes in fast, structured generation for JSON, regex, and agentic workflows using its RadixAttention prefix caching. It delivers significantly faster inference, especially for tasks with repeated prefixes, making it ideal for complex, structured outputs and multi-turn conversations. Choose SGLang over alternatives like vLLM when you need constrained decoding or are building applications with extensive prefix sharing.

View skill