Extract Lyrics from Song
Upload an MP3, WAV, or music video and turn sung vocals into editable plain text. This workspace lets you extract lyrics from song files without a separate converter — up to 200 MB, 99+ languages.
The editor on this page is the product, not a teaser. Drop a recording, run speech recognition on the vocal line, then copy or download a .txt draft you can paste into notes, karaoke prep, or a writing session. You stay on the same URL from upload to export. Treat the result as a first pass: verses and choruses usually land well on a clear mix; names, ad-libs, and stacked harmonies still need a human listen.
Drop your audio or music video here
or click to browse — MP3, MP4, WAV, and more
MP3 · MP4 · WAV · M4A · AAC · MOV · MKV · WEBM · OGG · FLAC (max 200 MB)
Try a demo
Public-domain vocal samples — click to load and transcribe
How to extract lyrics from audio
Upload your song file
Drop an MP3, WAV, M4A, or a music video with a vocal track. Common audio and video types are accepted up to 200 MB. If the file is a clip with an embedded soundtrack, the audio is taken automatically so you do not need a desktop converter first.
AI transcribes the vocals
Faster-Whisper large-v3 listens for sung or spoken words and writes them as readable text. Leave language on auto-detect, or pick from 17+ options when you already know the language. Clear lead vocals in a moderate mix give the most reliable lines.
Copy or download lyrics
Copy the draft to the clipboard or download a .txt file. Paste it into a notes app, a lyric document, or a songwriting toolkit. This page does not emit SRT cues; use the Transcribe tool when you need timestamps.
What this lyrics extractor includes
Audio and music-video uploads
Bring MP3, WAV, M4A, AAC, MP4, MOV, MKV, and similar files — up to 200 MB per upload. A music video is fine: the soundtrack is pulled automatically so you can recover the vocal text without ripping the file yourself.
Speech recognition tuned for singing
The model is built for speech, including sung lines across 99+ languages. It can still read a lead vocal sitting in a full mix, but a dry vocal, a rehearsal take, or a quieter backing track is easier to decode than a wall of synths.
Plain-text output, not captions
You get a clean document without SRT timestamps — made for copying, editing, and study. If you need karaoke-style cues or burned-in subtitles, export is the wrong format here; switch to the transcription workflow for timed subtitles.
Private processing
Files are deleted after processing and are not stored as a public library. Copy the text or download .txt when you are done. Use this for recordings you have a right to work with; do not treat the export as a license to republish someone else’s catalog.
A practical guide to recovering song lyrics
People search for a way to recover words from a recording they already have: a demo, a rehearsal, a live take, or a video where the vocal is clear enough to hear. This page is speech-to-text for that job. It is not a licensed lyric database and it does not look up published sheets. You upload media, the model writes what it hears, and you leave with a draft you can edit. The notes below cover sources that work, how singing differs from podcast speech, what to check before you paste the text anywhere public, and where a different Pixwit tool is a better next click.
What to upload, and what to skip
A usable source has a lead vocal you can follow with your ear. Phone voice memos, rehearsal-room recordings, stems, and music videos with a prominent singer all work. Problems usually come from extreme compression, a crowd louder than the singer, or an instrumental with no words. Instrumental beds, lo-fi vinyl noise, and files under a few seconds of singing may return little or nothing — that is honest, not a bug. Keep uploads within 200 MB. Audio types include MP3, WAV, M4A, AAC, OGG, and FLAC; video includes MP4, MOV, MKV, and WebM. If you only want to see the pipeline, load a public-domain demo on this page, then swap in your own track.
Why sung lines are harder than spoken interviews
Singing stretches vowels, stacks harmonies, and often sits behind drums and bass. A speech model still maps those sounds to words, but it may merge backing vocals into the lead, miss a whispered bridge, or spell a made-up hook phonetically. Melisma (many notes on one syllable) can look like extra letters. Rap with heavy bounce and ad-libs is closer to speech and often transcribes more cleanly than a washed-out pop chorus. If you have a vocal stem, use it. If you only have a stereo mix, turn down competing noise at the source when you can, or try a shorter excerpt around the verse you care about.
Language, auto-detect, and mixed-language tracks
When you know the language, set it so the decoder is not guessing among similar vowels. Auto-detect is reasonable for a single clear language and weaker when the hook switches mid-line (English chorus, Spanish verse, and so on). Accents, artist names, and invented spellings are the first things to check. A K-pop title that mixes romanization with hangul, or a jazz standard with French lines, will need a pass from someone who knows the words. This workspace is not a translator: it writes in the language it hears. It also does not promise word-level karaoke timing or speaker labels for a duet.
How to judge a draft before you copy it
Play the first chorus against the text: did the title phrase, the rhyme at the line end, and the last word of the hook land? Skim for missing negations (not / never), doubled lines from a delay effect, and numbers in a lyric that change meaning if they are off. Repeated refrains may appear once or several times depending on how clearly they were sung. Empty output on a stretch of synth is expected. Do not force words onto a beat drop. If the file is a cover you recorded, you still own the performance; the text is a convenience, not a substitute for crediting the original writers when you share the song.
Privacy, limits, and tools this page is not
Processing can take a short while on a long file, and a timeout applies if the job stalls. Stay within 200 MB. Media is sent to the speech backend so the model can hear it; it is not published in a gallery. This editor will not isolate stems, will not pitch-correct, and will not generate a new vocal. For timestamped captions, use AI Transcription. For a video that needs burned-in lines, export text here only if you want a document — then move to a caption tool with SRT. Related links on this page stay inside audio and lyric workflows so the next click matches the job you started.
Background reading: Speech recognition (Wikipedia)
Who uses this lyrics workspace
Songwriters recovering a demo
You hummed a verse into a phone and need the words on a page. Speech-to-text is the first pass; you still fix the hook spelling before you send it to a collaborator.
Covers, karaoke, and rehearsal
A practice track or a live video is easier to learn from when the lines are on screen in a notes app. This export is plain text, not a licensed karaoke file and not a timed .lrc.
Teachers and language learners
A clear vocal in another language can become a study sheet. Confirm grammar and names against the recording. Do not treat an automated draft as a published translation.
Producers labeling a session
Engineers paste a rough lyric sheet into a session note so a singer can jump to a section. The text is a map, not a legal credit list — writers still belong in the metadata you ship.
Extract Lyrics from Song FAQ
Related tools
Run this lyrics extractor on this page
Upload a track in the editor above, review the draft, then copy or download when the lines look right.
