Set out the same words in the same order, and the listener's reception will differ entirely depending on who utters them and in what voice. The voice is not a mere sound but a medium that conveys the speaker's inner state (tension, ease, sincerity, haste) ahead of the speaker's intention. What this volume asks is how three elements of the voice — tone (pitch, timbre), pace (tempo), and "ma" (the pause, silence) — produce the impression of confidence. The precise meaning of Mehrabian's non-verbal research, Klofstad's research on the perception of leadership, the Apple-Streeter-Krauss speed experiment, Anne Karpf's anthropology of voice, Zeami's theory of vocal music, and Tatekawa Danshi's theory of "ma." Six traditions draw out a path to compose the voice as a technique.

01The Voice as "Another Self-Presentation" — What Arrives Before the Words

Day to day, we are usually conscious of the content of communication as words. What to say, how to assemble the logic, what evidence to bring. But on the listener's side, the voice itself arrives before the words can be decoded. Pitch, pace, tremor, pause — all convey the speaker's inner state in an instant, and that first impression is not easily overwritten by content that follows.

For the regulatory reviewer this is an especially important phenomenon. Meetings with the requester, the delivery of corrections, agreement with the originating client — in every scene, before words are chosen, the voice is already setting the temperature of trust or wariness. Alongside the "Outside → Inside" circuit treated in Vol. 3 (Cuddy's power posing, Bem's self-perception), the voice is the most important constituent element of the exterior. This volume composes that voice not from impressions but from the accumulated record of empirical research and classical writing.

02Mehrabian's "7-38-55" — Misunderstanding and Core

When we speak of non-verbal communication, a certain number is invariably cited. "In the transmission of a message: words 7%, voice 38%, expression 55%" — the so-called Mehrabian's rule. The figures put forward by the American psychologist Albert Mehrabian (1939–) in research from 1967 and 1971.

But these numbers are widely misunderstood in public discourse. This is not the claim that "in all conversation, only 7% comes through as words." Mehrabian himself has emphasised this repeatedly in later years. The original experiment yielded these figures under the limited condition of "what does the listener believe when the messages of words, voice, and expression contradict one another?"

"My research is limited to the communication of feelings or attitudes — particularly cases where the verbal content and the non-verbal cues are inconsistent. Applying 7%, 38%, 55% to all communication is not what I intended." — paraphrase of Albert Mehrabian (a note on his own website)

Even when correctly understood, Mehrabian's implication is strong. In scenes that convey feelings or attitudes, voice and expression surpass words. For example, if "It's fine" is delivered in a trembling voice with a stiff face, the listener concludes that "it is not fine." When a reviewer says "no problem" to a requester, if the voice is awkward, the requester believes the voice, not the content.

03Pitch Decides the Perception of Leadership — Klofstad's Research

The effect of voice pitch on the judgement of others was demonstrated by the American political scientist at the University of Miami, Casey Klofstad (1976–), in a large-scale experiment. His 2012 paper in Proceedings of the Royal Society B.

"We recorded the same utterance at manipulated higher and lower pitch and asked subjects which speaker they would choose as leader. For both men and women, there was a marked tendency to choose the lower-pitch voice as leader. This carries implications for biological evolutionary psychology, but also for social learning." — paraphrase of Klofstad, Anderson & Peters, Proc Roy Soc B (2012)

Klofstad's finding has two implications. First, a lower-pitched voice unconsciously carries the impression of "competence, confidence, trust." Second, because the voice tends to rise under tension, training to consciously lower the pitch at moments of tension produces the impression of confidence.

Caution is required, however. A voice forced unnaturally low becomes unnatural and counterproductive. Klofstad himself recommends "slightly lower" within a natural range. For the reviewer, the practical implication is to take a deep breath that opens the chest before a tense meeting with a requester — and recover the natural lower register.

04The Impression Brought by Pace — The Classical Apple-Streeter-Krauss Experiment

There is a classical study on the other element of the voice, pace. In 1979, the psychologist at Columbia University Robert Krauss and his collaborators (William Apple, Lynn Streeter) conducted an experiment. They let subjects listen to recordings at varied speaking pace and rate the speaker's impression.

"Fast speech carries the impression of intelligence, competence, and persuasiveness, but also the impression of being 'pushy' and 'lacking sincerity.' Slow speech carries the impression of sincerity and thoughtfulness, but also the impression of 'inadequate competence' and 'lack of confidence.' Pace is a double-edged sword, and its meaning reverses depending on context." — paraphrase of Apple, Streeter & Krauss, JPSP (1979)

The implication of this experiment is that it is not a simple matter of "faster is better" or "slower is better." Adjustment of pace to the context creates the impression of confidence. Concretely, the following applies.

Faster

To convey competence and persuasiveness

In the opening of a presentation, in agreement-building, in scenes that should convey energy — a slightly faster tempo is effective. But if too fast, it turns into the impression of haste or pushiness.

Slower

To convey sincerity and thoughtfulness

When delivering a hard finding, sharing a cautious judgement, or leaving room for the listener to think — a slightly slower tempo is effective. But if too slow, it turns into the impression of a lack of confidence.

Variation

Switch pace according to the scene

The thing most to be avoided is speaking at a constant pace. Slow at important points, normal pace for the flow of explanation — the very change of pace supports the listener's concentration and understanding.

05The Dynamics of "Ma" — What Silence Conveys

Of the three elements of voice, the one least consciously attended to, and the most powerful, is "ma" — that is, silence. Ma is not voice, but it is part of the structure of voice. Speech without ma becomes a monotonous stream of sound; speech with excessive ma makes the listener uneasy. The one who can place an appropriate ma is the one who can express confidence in the spaces of the voice.

The American linguist Deborah Tannen (1945–), in That's Not What I Meant! (1986), observed cultural and gender differences in the sense of ma. On the U.S. East Coast a short pause is preferred; on the West Coast and in Japan, longer pauses are tolerated. The same span of silence, depending on the listener's cultural background, is received either as "thinking" or as "stuck."

In the pharmaceutical workplace, especially in Japan, the use of ma is directly tied to the building of trust. When one receives a sharp question from a requester, placing a half-second to one-second ma before answering — rather than answering reflexively — carries the impression of thoughtfulness. This is not "pretending to think" but the technique of weaving real time-to-think into the structure of the voice.

06The Embodied Voice — Anne Karpf's Anthropology of Voice

The one who observed the voice at the intersection of physiology, anthropology, and sociology is the British journalist and sociologist Anne Karpf (1950–), in The Human Voice (2006). What she shows is the perspective that the voice is not a mere vibration of air but an expression of the body itself.

"The voice is not an abstract signal severed from the body. The diaphragm, the chest cavity, the vocal folds, the oral cavity, the nasal cavity — the cooperation of all bodily tissue makes the voice. When tension stiffens the chest, the voice becomes shallow. When the breath is shallow, the voice trembles. To change the quality of the voice, the body must be changed." — paraphrase of Anne Karpf, The Human Voice (2006)

Karpf's implication is practical. To "produce a voice of confidence" is, before being a question of technique, a question of the body. Deep breathing, opening of the chest, relaxation of the shoulders — without the preparation of the body, mere manipulation of the voice has its limits. This resonates with William James's body-first hypothesis and Cuddy's power posing treated in Vol. 3.

Karpf also notes that the causes of the contemporary shallowness of voice include sedentary living, shallow breathing, and screen-mediated communication. To consciously breathe deeply and to fold into daily life time in which the voice is supported by the whole body — this is at once a classical and a contemporary practice.

07Zeami's Theory of Vocal Music — The Law of Voice in Noh

From the Eastern lineage, the one who treated the technique of voice most systematically is Zeami (c. 1363 – c. 1443). As the master of Noh, he discussed in Kakyō (1424) and Fūshikaden (c. 1400) the discipline of voice the Noh actor must hold on stage. The term Zeami used is ongyoku (音曲) — the total art of voice and sound.

"Senbun kōken (先聞後見) — the audience first hears the voice and afterwards sees the actor's figure. The voice makes the temperature of the place; the figure reinforces it. The actor who cannot rule the place by voice, however beautiful the figure, cannot move the audience's heart." — paraphrase of Zeami, Kakyō (1424), section on vocal music

Zeami's senbun kōken points to the same place that contemporary science has confirmed after the fact. The voice arrives before the figure. What Mehrabian showed by experiment, what Klofstad confirmed by data — Zeami had already extracted the same proposition from stage practice six centuries earlier.

Zeami further set, as a requirement for the Noh actor, training to consciously change the register of the voice according to scene. The concepts of koe-tsuki and koe-awase are concrete techniques that decompose vocal expression and train it. The implication for the reviewer is clear: the voice is not a gift of birth but a technique that can be composed by training.

08The "Ma" of Rakugo — Learning the Breath of Speech from Tatekawa Danshi

There is a tradition of voice native to Japan: the ma of rakugo. Beichō, Shinchō, Danshi — the post-war masters honed ma as the technique of designing the audience's laughter. Known as a theorist among them, Tatekawa Danshi (1936–2011) discussed the essence of ma in Modern Rakugo Theory (Gendai Rakugo Ron) (1965).

"Ma sits in the centre of the art. Between voice and voice, between word and word, there is the margin in which the audience's heart moves. The performer is only 'speaking' for half the time. The other half is the time the audience is 'listening' — that is, the time ma creates. The performer who cannot hold ma cannot draw in the audience's heart." — paraphrase of Tatekawa Danshi, Modern Rakugo Theory (1965)

What Danshi shows is the observation that ma is not "time when one is not speaking" but the active time in which the listener's understanding and feeling move. Without ma, the listener becomes passive; with too much ma, the listener becomes uneasy. The performer who can place an appropriate ma has folded the movement of the listener's heart into the structure of the voice — a principle that applies, just as it stands, to the reviewer's dialogue and to the field of education.

The practical implication: before delivering an important finding, hold one beat. When you receive a question, do not answer immediately, hold half a beat. When you switch the flow of speech, take a clear ma. These are not staged "build-up" but the practice of making the margin in which the listener's heart moves.

09In the Pharmaceutical Workplace — The Role the Reviewer's Voice Plays

Let us bring the abstract down to the field. Scenes in the reviewer's daily work in which the three elements of voice (tone, pace, ma) are tested.

Scene 1

The moment of delivering a correction (pitch and pace)

The harder the finding, the more the voice tends to rise under tension. Consciously take a deep breath that opens the chest and return to the natural lower register. Pace slightly slow — leaving the requester room to think. Field application of the implications of Klofstad and Apple-Streeter-Krauss.

Scene 2

Immediately after a hard question (ma)

Rather than answering reflexively, hold a half-second to one-second ma, then answer. The ma carries the impression of "thinking" and conveys sincerity. The Danshi-like deployment of ma.

Scene 3

Scenes of agreement-building (variation of pace)

Drop the tempo at important points of agreement; return to normal tempo for supplements and explanations. Constant-pace utterance bores the listener — variation of pace supports concentration.

Scene 4

Extended explanation (body and voice)

The quality of the voice is tied directly to the state of the body. Karpf's implication — relax the shoulders, breathe deeply, support the voice with the whole chest. Without this, the voice in the afternoon is shallower than in the morning.

10Five Practices for Composing the Voice

To bring the same place to which the six traditions pointed down into daily practice, five practices are set here.

Practice 1

Three breaths before speaking (Karpf)

Immediately before an important utterance, take three deep breaths. Open the chest, lower the diaphragm, compose the bodily foundation of the voice. This alone naturally lowers the pitch and reduces the tremor.

Practice 2

Lower pitch within a natural range (Klofstad)

Consciously lower the slightly-too-high voice of tension. But do not force it unnaturally low. Awareness of chest resonance returns one without strain to the lower register.

Practice 3

Change pace by the scene (Apple-Streeter-Krauss)

Important findings and difficult explanations: slow. Flow and supplements: normal tempo. Scenes that convey energy: faster. Merely avoiding constant pace changes the listener's concentration dramatically.

Practice 4

Do not fear ma (Zeami, Danshi)

Do not bridge the silence with "uhm" or "ah." Half a second of silence is conveyed to the listener as time spent thinking. Make use of ma as the margin of the voice.

Practice 5

Do not generate contradiction (Mehrabian)

Align words, voice, and expression in the same direction. The voice that says "It's fine" should be the voice that genuinely is fine. Contradiction sends the listener the signal of "lie."

In Closing

The voice conveys to the listener far more than the speaker realises. Tone, pace, ma — three elements that, before the words are decoded, carry the impression of confidence and sincerity. The correctly understood Mehrabian, Klofstad's pitch research, the Apple-Streeter-Krauss speed experiment, Karpf's embodiment of voice, Zeami's senbun kōken, Danshi's philosophy of ma. Six centuries of Noh and twenty-first-century psychology have pointed, in different words, to the same place. To compose the voice is to compose the interior; to compose the interior is what shows in the voice. The circuit of body-first treated in Vol. 3 runs through the voice as well.

The greatest message of this volume is that the voice is not a gift of birth but a technique that can be composed by training. Deep breath, natural lower register, variation of pace by scene, the courage not to fear ma. These are not special gifts; through the accumulation of conscious practice, anyone can acquire them.

Key Points — Three to Take Away
  1. Mehrabian's "7-38-55" is a figure limited to cases where feeling or attitude contradicts the words, but the implication is strong — voice and expression can surpass words. If a reviewer's voice is stiff, the requester believes the voice, not the content.
  2. The three elements of voice (pitch, pace, ma) are techniques that can be composed by training. Klofstad showed a naturally lower pitch, Apple-Streeter-Krauss showed variation of pace by scene, Zeami and Danshi showed the margin of ma, each as a structure of confidence.
  3. The quality of voice is the quality of the body. As Karpf shows, without deep breathing, an open chest, and relaxed shoulders, mere manipulation of the voice has its limits. Three breaths before speaking is the minimal unit of practice.
References
  1. Mehrabian, Albert & Wiener, Morton. "Decoding of inconsistent communications." Journal of Personality and Social Psychology 6 (1), 1967, pp. 109–114.
  2. Mehrabian, Albert. Silent Messages: Implicit Communication of Emotions and Attitudes. Belmont, CA: Wadsworth, 1971.
  3. Klofstad, Casey A., Anderson, Rindy C. & Peters, Susan. "Sounds Like a Winner: Voice Pitch Influences Perception of Leadership Capacity in Both Men and Women." Proceedings of the Royal Society B: Biological Sciences 279 (1738), 2012, pp. 2698–2704.
  4. Apple, William, Streeter, Lynn A. & Krauss, Robert M. "Effects of pitch and speech rate on personal attributions." Journal of Personality and Social Psychology 37 (5), 1979, pp. 715–727.
  5. Scherer, Klaus R. "Vocal affect expression: A review and a model for future research." Psychological Bulletin 99 (2), 1986, pp. 143–165. (a survey of voice-and-emotion research)
  6. Karpf, Anne. The Human Voice: How This Extraordinary Instrument Reveals Essential Clues About Who We Are. London: Bloomsbury, 2006.
  7. Tannen, Deborah. That's Not What I Meant! How Conversational Style Makes or Breaks Relationships. New York: William Morrow, 1986. (cross-cultural study of "ma" in conversation)
  8. Zeami, Kakyō (Japan, 1424). Sections on vocal music and senbun kōken. (Annotated: Omote & Katō, eds., Zeami / Zenchiku, Nihon Shisō Taikei 24, Iwanami Shoten, 1974)
  9. Zeami, Fūshikaden (Japan, c. 1400). Chapter on vocal music. (same annotation)
  10. Tatekawa, Danshi. Gendai Rakugo Ron (Modern Rakugo Theory). Tokyo: San-Ichi Shobō, 1965. (the foundational work that systematised the theory of ma)
  11. Cicero, Marcus Tullius. De Oratore. 55 BCE, Book III. Discussion of voice (vox) and delivery (actio).