SHADOWWW · Voice Models
Four takes. Four voices. Free to create with.
RVC v2 · 48 kHz · 200 epochs · CC BY-NC 4.0
Each model here is a single session, captured as recorded. Shadowww cut four vocal takes, and each take became its own model: its own tone, its own register, its own amount of tuning. Nothing was pitched, tuned or polished after the fact. Only the silence was taken out.
Pick the one that fits your track, load it into any RVC v2 tool, and sing through it.
The voices
| Voice | Register | Tuning in the take | Best for | Download | |
|---|---|---|---|---|---|
| T02 | Raw and low | C♯3 · ~141 Hz | None, fully natural | Low, unprocessed, intimate vocals | zip |
| T03 | Bright and lifted | G3 · ~190 Hz | Moderate | Higher melodies, hooks | zip |
| T04 | Natural mid | F♯3 · ~181 Hz | None to light | Mid-range, conversational lines | zip |
| T05 | Polished mid | E3 · ~166 Hz | Strongest | A tighter, auto-tuned sound | zip |
"Tuning in the take" is measured, not a setting: the share of sung notes landing within 15 cents of a semitone. Natural singing sits around 30%: T02 22% · T03 45% · T04 28% · T05 51%.
Not sure which one? Start with T02 for a natural sound or T05 for a tuned one. T04 was trained on the least audio (3.2 min); it is great for short phrases and can waver on long held notes.
Alternates. T02 and T03 each include an earlier checkpoint (-e120, -e140) from the point
where training loss was lowest. If the final model sounds over-processed or metallic on your
material, try the alternate.
Quick start
- Download a voice's
.zipfrom the table above. It holds the model (.pth) and its index (.index). - Load it in your RVC v2 tool:
- Applio: open the Download tab, paste the zip link from the table, and download; the voice then appears in Inference.
- Replay (Mac): unzip, then put the
.pthand.indexin a new folder inside~/Library/Application Support/Replay/com.replay.Replay/models/and restart Replay. - RVC WebUI: put the
.pthinassets/weightsand the.indexinlogs/, then refresh.
- Feed it a clean vocal: dry, solo, no reverb, no backing track. Separate the vocal first if you are starting from a full song.
- Start from these settings and adjust by ear:
| Setting | Start at | What it does |
|---|---|---|
| Pitch extractor | rmvpe |
Most stable pitch tracking for singing |
| Transpose | 0 from a male voice, -12 from a female voice |
Moves the input into the model's range |
| Index rate | 0.6 (range 0.5–0.75) |
Higher = more Shadowww timbre, lower = clearer words |
| Protect | 0.33 |
Keeps breaths and consonants natural |
| Filter radius | 3 |
Smooths pitch jitter |
If it sounds off: robotic or warbly means lower the index rate or try the alternate checkpoint. Cracking on high notes means transpose down, or try T03. Muddy words mean lower the index rate to about 0.5 and make sure the input is dry.
What's in each folder
t02/ shadowww-t02.pth · added_shadowww-t02.index · shadowww-t02.zip
shadowww-t02-e120.pth · shadowww-t02-e120.zip · dataset.json
t03/ same layout, alternate = e140
t04/ shadowww-t04.pth · added_shadowww-t04.index · shadowww-t04.zip · dataset.json
t05/ same as t04
SHA256SUMS checksum for every file
dataset.json records how each take was prepared. No audio is published.
How they were made
| Framework | Applio 3.6.5 · RVC v2 |
| Sample rate | 48 kHz |
| Pitch extractor | RMVPE |
| Embedder | ContentVec |
| Pretrained base | KLM 4.1 48k (generator and discriminator) |
| Training | 200 epochs · batch 8 · 2 × NVIDIA T4 |
| Training audio | T02 5.9 min · T03 8.2 min · T04 3.2 min · T05 4.8 min |
Data preparation. Each take is a raw vocal recording by Shadowww. The only processing was silence trimming: every pause of 0.4 s or more was removed, using a quiet threshold set per take at the split between room noise and voice, with 0.15 s kept at each edge. Audio was then peak normalised to −1 dBFS. No pitch shifting, tuning, effects or vocal separation were added; any tuning you hear was part of the original performance.
Licence and fair use
Released under CC BY-NC 4.0.
- ✅ Free to use, remix and share, with credit: "Voice model: Shadowww"
- ✅ Covers, demos, experiments and personal projects
- ❌ No commercial use without permission from Shadowww
- ❌ No impersonation: don't present AI output as a genuine Shadowww recording, and don't use the voice to deceive, harass or defraud anyone
- 🏷️ Label AI-generated vocals as AI
The voice belongs to Shadowww, who recorded, trained and published these models.
Credits
- Voice and recordings: Shadowww
- Pretrained base: KLM 4.1 by SeungHyun Han (Seoul Streaming Station), SeoulStreamingStation/KLM4.1, used under its licence, which requires this acknowledgement
- Training framework: Applio
SHADOWWW · three w's, one voice, four takes