Are there words in my head?
Peter Langland-Hassan
University of Cincinnati
When we say something to ourselves in inner speech, do we produce words? We certainly talk as though we do. How can we say something in inner speech, if we produce no words? Where are the words we bring into being? Where else could they be than within the skull of the person doing the inner speaking?
As natural as these inferences may be, there is a quick way to cast doubt on them. We can start by granting that, when we have inner speech, we are aware of words (in some sense of ‘aware’). But we are not simply aware of words. We are aware of them in a certain way, similar to the way in which we are aware of words when we hear them spoken. Specifically, we are aware of their sonic and phonemic properties, which account for the way the words sound. This is why we can use inner speech to judge whether two visually dissimilar words—‘lie’ and ‘fly’, or ‘floor’ and ‘oar’—rhyme. And it’s why we can say the same set of words in inner speech with many different pitches, accents, and intonations, in each case changing the sonic and phonemic properties we are aware of.
Now connect this observation to the thought that we have a general capacity to imagine sounds. Typically, when we imagine the sound of something, be it a car horn or plucked guitar string, neither the thing making the sound nor the sound itself is present within the mind of the person doing the imagining. We represent the sound of a car horn—and in that sense are aware of a certain sound—without there being a horn or horn sound thereby brought into existence. Applying the same reasoning to the words we are aware of in having inner speech, this suggests that they, too, fail to come into existence simply in the act of our representing them. The earlier temptation to say that there are words in the head when we generate inner speech simply mistook representing an x—and being aware of it in that sense—for being aware of an actual x present within the mind.
There are different reactions one might have to this clash of perspectives. Here I will meditate on just one. It is what I will call the double-duty response, as it looks to have things both ways: inner speech episodes do represent words, it allows, by representing their sounds (just as we might represent a tomato by representing its shape and color). Yet this is compatible with an inner speech episode also being an instance of the very words represented. So, whereas the representation of a car horn’s sound is not itself a car horn, something in the nature of inner speech (and, perhaps, of words) allows it simultaneously to play both roles: to represent a set of words and to be an instance of the words represented. It is a satisfying reconciliation if it can be made to work.
Of course, the double duty response owes an account of why we should think this double-duty scenario holds in the case of inner speech. An ingenious proposal, due to Daniel Gregory and Daniel Stoljar (and recounted in Gregory (2026)), notes that a photograph of Los Angeles’s famous ‘Hollywood’ sign both represents the word ‘Hollywood’ and—if we isolate the relevant parts of the image—contains an instance of the word ‘Hollywood.’ The example suggests—plausibly, I think—that images of words can indeed both represent words and contain instances of the words represented. No such thing seems possible for a car horn.
However, not all images of words themselves contain instances of the words represented. Imagine taking a photograph of the Hollywood sign from directly above it. In that case, the part of the image representing the sign would look roughly like nine horizontal bars—one for each letter, as seen from above. While the ‘Hollywood’ of the Hollywood sign is still represented by the photo, the photo does not contain an instance of the word ‘Hollywood’. Apparently, only some images of words contain instances of the words they represent.
What feature allows double-duty representation to occur in some images of words? In the first Hollywood sign example, both the sign pictured and the parts of the photograph that represent the sign instantiate the canonical shapes of the letters associated with the word ‘Hollywood’. This is missing when the sign is photographed from above. Following that logic, for an ordinary episode of inner speech to both represent an utterance of ‘It’s hot’ (via representing its sonic and phonemic properties) and to be an instance of the words ‘It’s hot,’ the inner speech episode itself would need to instantiate the canonical sonic and phonemic properties associated with overt utterances of ‘It’s hot.’ But, if that is so, the analogy falls apart, as inner speech is strictly speaking silent.
However, the defender of the double-duty view might venture the following reply: for an inner speech episode of my saying “It’s hot” to also be an instance of the words ‘It’s hot,’it is enough that there is a structural isomorphism between the inner speech episode and the words it represents,such that there is a systematic mapping of features of the wordsrepresented onto features of the inner speech episode itself. On this view, a representation of the sonic and phonemic properties of an utterance of “It’s hot” needn’t itself make a sound in order to be an instance of the words ‘It’s hot’; it need only have structural features that map in a systematic way to sonic features of the overt utterance ‘It’s hot,’ such that the former can be recovered from the latter through pattern-matching.
By invoking the notion of a structural isomorphism between representation and that which is represented, this double-duty view suggests that conditions commonly held for being an analog representation hold for inner speech (see, e.g., Kulvicki (2014), Beck (2018), and Maley (2011) on the nature of analog representation). Whether or not inner speech episodes are correctly viewed as analog in this sense, the possibility reveals a weakness in the double-duty theorist’s response. An analog representation of an x is not normally itself an x. Consider the standard examples of analog representation: the height of mercury in a thermometer, a speedometer dial, the grooves on a vinyl record. In each case, there is a systematic covariation between properties of the representation and the properties represented (temperature, speed, and sounds, respectively). But in no case are we inclined to hold that the representation is itself an instance of what is represented. So, it remains unclear why viewing inner speech episodes as structural analogs of the utterances they represent would vindicate the double-duty view.
An alternative defense of the double-duty view rejects the idea of a structural resemblance between representation and what’s represented and holds instead that what makes an inner speech episode an instance of the words it represents are facts about how the relevant mental representations are used. I cannot adequately address that possibility here. But I end with a note of caution: we should not confuse a person’s acting in similar ways when they speak and when they imagine speaking with their using words in each case. It may be enough that they put themselves in similar mental states each time. I can, for example, hone my ability to catch frisbees both by catching frisbees and by imagining catching them. There is nothing in the latter case that gets tossed like a frisbee.
References:
Beck, J. (2018). Analog mental representation. WIREs Cognitive Science, 9(6), e1479. https://doi.org/10.1002/wcs.1479
Gregory, D. (2026). Inner speech and sign languages. Synthese, 207(4), 176. https://doi.org/10.1007/s11229-026-05554-5
Kulvicki, J. V. (2014). Images (First edition). Routledge, Taylor & Francis Group.
Maley, C. J. (2011). Analog and digital, continuous and discrete. Philosophical Studies, 155(1), 117–131. https://doi.org/10.1007/s11098-010-9562-8
So what I found interesting is that when we learn psychology and perception, cognition and language – we learn about inner thoughts, thought processes. These are not really physical things but exist more as a psychological symbolic structural forms to denote our brain activity.
The concept of inner speech can be seen through the view of a visual type learner in the same way as for an auditory type learner? Could our primary abilities, skill set and characteristics define our inner speech? Is that why some people would rather write down a thought, a question or an idea to form that representation?
“when we imagine the sound of something, be it a car horn or plucked guitar string, neither the thing making the sound nor the sound itself is present within the mind of the person doing the imagining”. I don’t think this is true (at least if by “sound” we mean “auditory experience”). If I imagine a car horn, either I’m inducing an actual (albeit faint) auditory experience, or I’m not really imagining the sound or a car horn at all and instead merely thinking about imagining it.