YouTube transcripts

Every Type of Vocal Stack Explained: video thumbnail

Every Type of Vocal Stack Explained transcript

Jaco Explains · @JacoExplains

Published September 10, 202614:15168.3K views

Watch this video on YouTube

Transcript analysisComputed from the caption text

Words

2,242

Runtime

14:15

Speaking pace

157wpm

Reading time

9min

157 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.

Opening (first 30 seconds)

A vocal stack is what happens when multiple vocal layers are recorded on top of each other to create a sound that a single voice cannot produce on its own. Most vocal stacking falls into recognizable categories, and understanding them changes how you hear recorded music permanently. Today, we're breaking down every type of vocal stack, what each one does, and why producers use them. Starting with the double track. The double track is the simplest vocal stack. >>

79 words, the words spoken in the first 30 seconds at 157 words per minute.

Sentence shape

MeasureThis transcript
Sentences176
Average words per sentence12.7
Longest sentence41 words
Questions asked2
Sentences containing a number11

Most used terms

  • vocal36
  • different21
  • stack19
  • harmony17
  • singer17
  • singing14
  • sound14
  • music13
  • melody12
  • lead10
  • voice10
  • recording9

Filler phrases

10 in total: like 8 · actually 1 · you know 1.

A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.

What this transcript is

Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.

Transcript

A vocal stack is what happens when multiple vocal layers are recorded on top of each other to create a sound that a single voice cannot produce on its own. Most vocal stacking falls into recognizable categories, and understanding them changes how you hear recorded music permanently. Today, we're breaking down every type of vocal stack, what each one does, and why producers use them. Starting with the double track. The double track is the simplest vocal stack. >> Don't you know that two is better [music] than one? >> And the starting point for almost every vocal arrangement in commercial music.

The singer records the lead vocal, then they sing the same part again. But no human being is able to sing the same line exactly the same way twice. One consonant arrives a few milliseconds earlier. One note sits a few cents sharper. A vowel lasts slightly longer. When you play those two recordings together, those tiny differences create thickness. Instead of hearing two distinct singers, you hear one voice that feels slightly larger.

This is why copying a vocal file and playing the exact duplicate underneath it does not work. If the waveform is identical, you've made the same recording louder. The double track effect comes from the differences, not from the repetition. John Lennon is known for using it. >> He blew his mind out in a [music and singing] car. >> However, he did not invent the technique. It was being used long before the Beatles. However, he did use it constantly and complained extensively about having to sing everything twice.

Abbey Road engineer Ken Townsend eventually developed artificial double tracking, or ADT, which used tape machines to create a slightly delayed and modulated version of the original recording. That gave Lennon something resembling a second performance without requiring him to actually sing the song a second time. Which tells you something important. The imperfections are the effect. If the second vocal becomes too different, the listener hears two separate performances.

Close enough and the brain fuses them. Producers can also decide how the doubles sit in the stereo field. Both centered, one slightly outward, or two doubles panned hard left and right. The unison stack. A unison stack takes the double track further. Instead of singing the melody twice, the singer records it three, four, or more times. >> [music and singing] >> There's still no harmonies yet. Imagine one lead vocal in the center, another performance on the left, another on the right, two more underneath.

Suddenly, the same melody occupies a much larger portion of the stereo field. This is one reason modern pop choruses can sound dramatically larger than the verses, even when the underlying instrumentation is not changed much. In the verse, there's usually one relatively exposed lead vocal. In the chorus, six versions of the same singer surround you. The vocal density can make a huge difference. There's also a point of diminishing returns.

When you stack enough performances and the individual imperfections begin averaging together, the sound becomes increasingly smooth and ensemble-like. Four closely recorded takes might still sound unmistakably like one singer. 20 takes can begin sounding more like a small crowd made entirely from that singer. The producer decides how tight or loose the stack should be. Very tight timing produces a polished modern sound.

Slightly looser produces something broader and more organic. Too loose and it stops sounding powerful and starts sounding messy. Vocal stacking is not simply recording the same thing many times. It's deciding exactly how much human variation to preserve. Harmony stacks. A harmony stack changes the notes. One man will go. vocal parts sing different notes around the lead melody rather than doubling it. The simplest example is a singer holding one note while another voice sings a third above it.

Put them together and you have harmony. Actual harmony arrangements are more complicated than just singing everything a third higher. The correct interval changes depending on the chord underneath, the melody, the range of the singers, and the character of the arranger wants. Sometimes the harmony sits above the lead, sometimes below, sometimes both. And then each harmony part can itself be doubled. A lead vocal in the center, a harmony above it recorded four times, another harmony below it recorded four times.

That's nine vocal recordings before a single ad-lib has been added. This is where stacks become enormous. Beyoncé's engineer, Stuart White, has described recording background vocals by stacking individual harmony parts four times with different groups organized separately in the session. One person can build a choir one part at a time. Queen did this decades earlier. Freddie Mercury, Brian May, and Roger Taylor repeatedly overdubbed themselves to create the vocal sections the band became known for.

On Bohemian Rhapsody, particularly through the operatic middle section, >> Thunderbolt and lightning, [music] very, very frightening me. >> the vocals are a construction built from repeated performances by three voices. Contemporary accounts describe extensive overdubbing and track bouncing because the available recording technology had a fraction of the tracks a modern digital session contains. The real power of a harmony stack is that you're not making the singer louder.

You're building harmony vertically. A single melody becomes a chord. Each note in that chord becomes several voices. This is how three people can sound like 30. The breath and whisper layer. >> Love is endless, [singing] don't [music] be >> This one works in an almost opposite direction. Instead of making vocals sound enormous, you make it feel close and intimate. A breath layer is a very quiet airy performance placed underneath or around the main vocal.

Sometimes it's sung with a breathy tone. Sometimes it moves toward a genuine whisper. Those are technically different things. A true whisper does not have the vibrating vocal fold fundamental frequency that gives a sung note its pitch. Breathy singing is still voiced and can follow the melody normally. In production, both are used for the same reason, texture. A quiet breathy layer contains detail, air moving, consonants, lip sounds, the beginning and ending of words, because it's usually recorded very close to the microphone.

Those details feel almost unnaturally intimate. You don't necessarily hear a separate vocal track, you just feel that the singer is closer. Billie Eilish probably jumps to most of your minds. >> [music and singing] >> But productions are more complicated than simply placing one whisper track underneath everything. Finneas has described building songs from large numbers of vocal layers, lead doubles, harmony stacks, background vocals, and processed ad-libs.

On "When the Party's Over", he described the session as containing roughly 100 vocal tracks. What sounds like a simple intimate vocal is often an enormous construction. More tracks don't always make something sound bigger. Sometimes they make it sound closer. The octave stack. The octave stack is technically a type of harmony stack, but deserves its own category because it sounds so different. Instead of adding a third or a fifth above the melody, you sing the same melody one octave higher or lower.

The note has the same letter name. A C becomes another C in a different register. Because of that, ears tend to fuse octaves together more easily than other intervals. An octave below the lead adds weight and darkness without introducing a new melodic shape. An octave above adds brightness and height and sometimes tension if the higher voice sits in a more demanding part of the singer's range. Combine both and the same melody now occupies three registers simultaneously.

This is particularly useful when a chorus needs to become dramatically larger without a complicated harmony part distracting from the main melody. The tune stays the same. Octave stacks are also useful when singers have very different natural ranges. A male and female voice can sing the same melodic shape in different octaves and create something that feels unified while preserving the contrast between them. >> [music] [singing] >> The important distinction is that an octave is still a harmonic interval.

Producers treat it separately because it functions differently in practice, not because it's a categorically different technique. The ad-lib stack. Ad-libs are not following the lead melody. They are >> everything happening around it. >> [singing] >> The runs, the oohs, the small responses at the end of lines, the high note that appears behind the final chorus, the phrase that surfaces once and disappears. In R&B, soul and gospel especially, these can become a second vocal arrangement sitting on top of the first.

Unlike a traditional harmony stack, the parts do not have to move together. One ad-lib might climb upward while another falls. One might hold the note while the lead keeps moving. Another might answer something the main vocal sang 2 seconds earlier. The result might feel like one singer having a conversation with several versions of themselves. Mariah Carey is one of the clearest examples. >> [music] >> She's spoken about stacking background vocals and about arrangements where melody parts and ad-lib parts occupy different spaces.

Vocal stacking is not something producers added around her performances later. It's part of how she thinks about vocal arrangement. This is why an ad-libs act should not be thought of as improvised material left in by accident. Several performances might be recorded, some phrases kept, others discarded. Certain ad-libs moved outward in the stereo field. A high run appearing only once because repeating it would reduce its impact.

Something designed to sound spontaneous can be one of the most carefully constructed parts of the record. The studio is being used as part of the composition. The final arrangement may be physically impossible for one singer to reproduce live because the recording contains the singer doing five things simultaneously. The choir stack. Most of the techniques above can be created by one singer recording themselves repeatedly.

A choir stack introduces something overdubbing cannot fully replicate, different people. No two voices are acoustically identical. Different singers have different vocal track shapes, different formants, different vibratos, different timing tendencies, and different timbres. Record one singer 20 times and you have 20 performances of one voice. Record 20 different singers and you have 20 different instruments. Even singing the same note, the differences between voices create a complexity that copies of the same voice cannot produce.

This is why an actual choir does not sound like one singer duplicated 20 times. >> [music and singing] >> Phil Spector's wall of sound productions in the early 1960s used large groups of session singers, sometimes 10 or more voices on a single part, to create a density that overdubbing one singer could not match. The technique was expensive and time-consuming. It produced results that cannot be perfectly faked by pressing copy and paste.

Combine choir stacking with individual overdubbing and the numbers become significant quickly. Three singers, each singing three harmony parts, each part gets doubled, 18 tracks. What matters is that the variation is no longer just variation between performances. It's variation between actual voices. That gives a sound a complexity that cannot be manufactured from a single source, regardless of how many times it's duplicated.

The processed stack. The processed stack takes vocals and alters them electronically before combining them. Pitch shifting, formant shifting, distortion, delay, heavy reverb, artificial harmonization, time manipulation, Auto-Tune used aggressively enough that you're meant to hear it. At this point, the human voice stops being a singer being recorded and becomes raw material for an instrument. Imogen Heap's Hide and Seek is the clearest example of a song built almost entirely from processed vocal stacks. >> Where [music] are we? >> The effect came from Heap singing through a Digitech vocalist harmonizer controlled from a keyboard.

The harmonizer generated additional vocal pitches according to the notes she played, allowing one live voice to become an electronically generated harmony structure. The accompanying voices were constructed technologically around a single human performance. This is different from Queen physically recording harmonies. Both create a vocal stack. They reach it from opposite directions. Then there is Auto-Tune used as an aesthetic rather than a corrective tool.

Kanye West's 808s and Heartbreak became one of the defining examples. >> Say you want to get [singing and music] me back and you going to show me. Say you walk around like you don't know me. >> Engineer and producer Mike Dean has referred to the heavily processed treatment associated with that era as the heartbreak sound. And that deliberately artificial vocal became significantly influential on the direction of pop and hip-hop that followed.

Modern production lets you push this as far as you want. Record one line, duplicate it, pitch one copy upward, shift the formants on another, distort a third, put a fourth through a large reverb, pan them separately around the original. Five layers, all beginning as the same human voice, now occupying completely different roles in the arrangement. That's the biggest idea behind vocal stacking in general. A vocal recording is not necessarily the finished vocal.

It's material. The producer can duplicate it, surround it, harmonize it, contrast it, or combine it with entirely new performances. The final result might contain two voices, or 10, or 100. If it's been done well, most listeners never consciously notice any of them. They just hear one vocal that somehow sounds wider, warmer, closer, bigger, or completely impossible.

The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.

Use this transcript

Three free tools that work on the material around a video like this one. No signup, no login.