Schenker’s ideas found their way into modern music linguistics via the most popular dissident and the most active conductor and composer in America, Noam Chomsky and Leonard Bernstein. Noam Chomsky developed a linguistic theory at the Massachusetts Institute of Technology in the 1950s and 1960s that was to shake the foundations of behaviorism, which was generally prevalent in America at the time. He had already laid the foundations for this in his dissertation in 1955.
Chomsky grew up in a linguistic environment dominated by the epoch-making work Methods in Structural Linguistics (1951) by his academic teacher Zellig Harris. The son of Jewish-Russian parents with roots in Ukraine combined American descriptivism, which limits linguistic research to the systematic description of external linguistic forms, with mathematical rules and elements of set theory. From the standard studies of American Indian languages, he developed a complex and abstract theory of the structural organization of languages, which was strictly limited to the description of external structures – hence the name «structuralism,» which describes this school of thought.
Chomsky leads research out of the emerging dead end of ignoring important semantic aspects by reframing questions about language comprehension and linguistic competence. He attempts to show that a competent speaker fundamentally knows more about a language than can be deduced from an analysis of the kind proposed by Harris.
Chomsky’s illustrative example from his book Syntactic Structures, the publication of his dissertation from 1957, is as follows. Consider the three sentences:
They all have the same surface structure, namely:
Obviously, the first sentence can be interpreted in two different ways: On the one hand, it can mean that the hunters are shooting. On the other hand, it is also possible to interpret it as meaning that the hunters are being shot at, which is a somewhat frivolous interpretation, but also grammatically correct in German. The so-called syntactic deep structure is different in this case:
Later, in Aspects of the Theory of Syntax (1965), Chomsky asks the simple question of how it is possible for a human being to acquire such a complex skill as language in the short period of early childhood. In the background are considerable doubts about the behaviorist doctrine that even intellectual achievements can be traced back to a stimulus/response pattern and the acquisition of knowledge through learning. This doctrine does not seem to explain the speed with which a child acquires a language and is able to generate a theoretically infinite number of meaningful sentences and make grammatical utterances that it has never heard before. Chomsky concludes that humans must have an innate language ability that is independent of the language they learn as children. He seeks to describe this innate language-independent ability in the form of a universal grammar.
Since Chomsky’s grammar emphasizes the creation or generation of sentences and seeks to shed light on this process of creation, it has been given the suffix «generative.» A common abbreviation for the theory is gTG for «generative transformational grammar.» gTG clearly distinguishes itself from its predecessor, structuralism, which emphasizes the external description of sentences. The system is called transformational grammar because it has rules that can be used to convert the deep structures that are independent of a single language into surface structures for a specific language, such as German, English, or Chinese.
The result of the search for a universal schema for the syntactic aspect of language is the «x-bar schema,» which is supposed to be valid for all languages. All building blocks of a sentence, whether they are nominal phrases, verbal phrases, or prepositional phrases, have a so-called head and extensions: For example, the head of the noun phrase «the Viennese composer Schubert» is the expression «Schubert.» This means that the noun phrase is split into two elements, «the Viennese» and «composer,» and the extensions are in turn split into further elements («the Viennese» and «composer»). This results in a tree structure for syntactic analysis with branching points where exactly two branches always start.
The sentence «The Viennese composer Schubert wrote numerous German dances» can be represented as follows based on the x-bar scheme:

The idea of genetically determined language competence has been vehemently questioned, for understandable reasons, primarily by behaviorist linguists; it disregards their basic premise, namely the conviction that humans are born as a «blank slate» and can acquire any skill through conditioning. (Modern popular proponents of this theory are, appropriately enough, fathers of several daughters: by deliberately training his daughters Sofia, Judith, and Susan to play chess, Father Polgar wanted to prove that anything is possible with careful education—and he succeeded: Judith and Susan topped the rankings of the world’s best female chess players. Father Williams, in turn, planned the same for his daughters Serena and Venus in women’s tennis. The two have long dominated the WTA world rankings at will. Despite the undisputed and spectacular successes in chess and tennis, however, there is little doubt today that humans also have innate abilities.
Chomsky’s theories continue to exert a strong fascination in all circles concerned with language and symbolic systems. From the perspective of music theory, as already mentioned, a certain external similarity between grammar trees and the x-bar scheme and Schenker’s analyses is unmistakable. It was therefore only a matter of time before the first attempts were made to transfer the results of linguistic research in the field of transformational grammar to music, in the hope of finally being able to pin down the linguistic or grammatical properties of music.
The prelude to this took place in 1973 in a series of six lectures at Harvard University entitled «Music: The Open Question» – under memorable circumstances. The speaker was none other than the charismatic composer and conductor Leonard Bernstein, creator of «West Side Story.» In the foreword to the printed version of the lectures, he describes the audience at the time as a mixture of «students, people without higher education, the security guard from next door, my mother, music experts with no understanding of linguistics, their counterparts, scientists with no interest in poetry, their counterparts.» He wants to take them all on a «journey to Chomsky Land.»
Right in the first lecture, he sets the course:
Year after year, linguists’ hypothesis that there is an innate grammatical ability (as Chomsky calls it), a genetically controlled language faculty that is universal, is becoming increasingly substantiated. This is an ability that is unique to humans; it testifies to the unique power of the human mind.
Well, music does nothing else.
But how do we explore musical universality with such a scientific tool as linguistic analogy? Music, it is said, is a process of rewriting, a kind of mysterious sensualization of our innermost feelings. Music is easier to describe in dazzling prose than in equations. Even a scientist as eminent as Einstein said: «The most beautiful thing we can experience is the mysterious.» Why, then, do so many of us constantly try to explain the beauty of music, seemingly robbing it of its mystery? In reality, music is not only a mysterious and descriptive art, it is also a child of science. It consists of mathematically measurable components: vibration numbers and durations, damping units and intervals. Therefore, any attempt to explain music must consist of a combination of mathematics and aesthetics, just as linguistics combines mathematics with philosophy, sociology or whatever else. It is precisely for this interdisciplinary reason that I am so fascinated by the new linguistics: it provides a new access to music. So why not conduct research in the field of musical language, just as there are already fields of research in linguistic psychology and sociolinguistics?
What fascinates Bernstein is the prospect of finding a universal grammar of music and thus being able to build the much sought-after bridge between the different musical cultures of the world: the realization of the age-old dream of the unity of mankind suddenly seems within reach in the realm of the spiritual experience of music, thanks to Chomsky’s work, which the exuberant Bernstein, in his search for what unites humanity, attributes to a «Beethovenian impulse,» ironically elevating the astute and delicate intellectual to the status of a romantic genius of emotion.
However, the musical jack-of-all-trades has less luck in implementing his thesis. For example, he develops parallels to transformation rules in the following way: the sentence «Jack loves Jill» is represented by an ascending major triad. Bernstein derives the negation «Jack does not love Jill» by transforming the major triad into a minor triad – after all, it is sad that Jack does not love the likeable Jill. The wild mixture of syntactic structures, the equation of an abstract operator («doesn’t») with a concrete situation, and the emotional reaction of a listener to a sentence seems touchingly naive. To add to the confusion, Bernstein neatly correlates the first note with «Jack,» the second with «loves,» and the third with «Jill (not),» thereby equating the non-operator with «Jill» and placing the minor third on the verb «loves.» The whole thing seems like a kind of linguistic cubism, like Picasso’s portraits in which a nose is painted from the side, an eye from above, and the mouth from behind. Presumably, Bernstein wasn’t taking the analogies all that seriously.
His analogies to the rules of transformation are more serious, especially because the development of a musical counterpart does not require recourse meanings, but can be limited to playing with forms (Bernstein refers to a diagram that looks something like this):
On the lowest rung (A), the basic musical elements are available to the creative will: pitches; keys with their own scales and chords; time measures with all their motor consequences such as tempos, etc. From these, certain combinations arise (B): melodic motifs and phrases, chord progressions, rhythmic figures, and so on. They are the musical «basic structure,» on the same rung as the basic structure of the language ladder. This basic structure can be transformed by rearrangement and permutation into what we have called musical prose (C). And here the two parts of the double ladder no longer correspond: As you can see, a prose sentence in language is already a surface structure, while musical prose is only conceivable as a deep structure. But we were prepared for that. We knew from the outset that music can never be prose. So let us take flight, and our process of re-transformation will produce the aesthetic surface that we call music.
Bernstein thus sees the difference between language and music in the fact that in the former, prose only begins at the surface structure, while in music it is a property of the deep structure.
In the second part of the second lecture (entitled «Musical Composition»), Bernstein illustrates his ideas with an example that was to become the symbol of all subsequent theories of the Chomsky school of music: the beginning of Mozart’s Symphony in G minor, K. 550.
Bernstein analyzes the beginning of the first movement and concludes that the first heavy beat is not found until the first beat of the third bar: Bernstein also subjects the rest of the movement to meticulous examination. It turns out that in bar 10, there is a conflict between interpretations of a strong and a weak beat. Bernstein sees all these analyses as an application of transformation rules. His conclusion:
It should now be clear to everyone that the surface structure we have examined stands in lively contradiction to the unbearably boring symphonic deep structure that I have assumed to be its basis; and that this conflict gives rise to a syntactic ambiguity in Mozart’s musical phrasing that results from asymmetrical interventions but is completely controlled, classically kept in check by the balance of proportions in Mozart’s sonata form.
Whether Bernstein’s results are sound or not is not decisive. What is important is that he has pointed out to others that Chomsky’s theories may be useful for a deeper understanding of the syntactic characteristics of music.
Through a seminar inspired by Bernstein’s lectures, two scholars – still students at the time – began to explore the connection between transformational grammar and music. Ten years later (1983), they presented one of the modern classics of 20th-century American music aesthetics «A Generative Theory of Tonal Music», or GTTM for short.
One of the two, Fred Lerdahl, is a composer by training with a penchant for theory and well versed in formal systems. The other, Ray Jackendoff, is a linguist and accomplished amateur musician. The preface to GTTM begins as follows:
In the fall of 1973, Leonard Bernstein gave the Charles Eliot Norton Lectures at Harvard University. Inspired by the insights that transformational-generative («Chomskyan») linguistics had provided into the structure of language, he argued for a search for a «musical grammar» that could explain human musical abilities. As a result of the lectures, numerous people in the Boston area became increasingly interested in the idea of linking music theory and linguistics. Subsequently, Irving Singer and David Epstein launched a seminar on music, linguistics, and aesthetics at Harvard University in the fall of 1974.
This seminar laid the foundation for the collaboration between Lerdahl and Jackendoff. The result of their relentless dialogue is a text that addresses the topic they set themselves with unusual precision.
In addition to generative transformational grammar, Lerdahl and Jackendoff draw on a second important theoretical framework: cognitive psychology. Their starting point for analyzing music is the way it is represented in the mind of a competent listener. In doing so, they seek to make statements about how the human brain processes music. The book attempts to model what is referred to in English as the «musical mind.” It could be translated as the «musical spirit» or the «musical brain.» Unfortunately, there is no such general word for “mind» in German, and translations such as «Geist,» «Vermögen,» «Sinn,» «Seele,» or «Repräsentation im Kopf» already evoke associations with existing theories that the general English term does not imply. Lerdahl and Jackendoff assume that the processing brain has properties such as modularity, specialization, and automation, and that the actual nature of musical processing is not fully accessible to consciousness. The authors also believe that their theory can be tested using experimental psychology (aspects of it have actually been empirically investigated with encouraging results, among others by the Belgian music psychologist Irène Deliège, who plays an important role in the development of cognitive music psychology in Europe).
According to some cognitive theories, the human brain solves problems by breaking them down into several sub-problems and processing them in parallel and, to a certain extent, independently of each other. GTTM follows this approach and forges a wide variety of «comprehension strategies» such as harmony, voice leading, gestalt perception, and so on into an integrated model of music interpretation—with the aim of also saying something about the nature of the human mind.
Lerdahl and Jackendoff’s draft of a generative theory of tonal music is based on cognitive diversity and is divided into four parts:
For all four areas, GTTM first establishes rules that determine which structures are «well-formed,» as the model theorist says. They merely determine which group structures are correctly formed and which are not, without saying anything about how the structures are arrived at.
«Decent» group structures, for example, divide pieces of music completely into smaller parts that must not overlap. A well-formed group structure looks something like this:

The rules of well-formedness for group structures, on the other hand, exclude things like this:

GTTM has an innovative feature: the theory goes beyond the type of rules that have been presented in connection with grammars. The rules of web grammar described above, as well as all those in classical logical systems without exception, are unambiguous. This unambiguity is precisely what makes them so powerful. It allows conclusions to be drawn without a shadow of a doubt or structures to be constructed so precisely that no ambiguities remain. They obey forms such as: «If A and B are present, choose B» or «In the grid, go first to the left, then to the right, then back to the left.» Such formulations leave no room for doubt and offer no scope for interpretation.
This is not the case with the second type of GTTM rule. These rules are more like more or less weighty hints or recommendations of the type: «If A and B are present and there are no other reasons to the contrary, then prefer A» or «In the grid, if it is easy to do so, go left first and then right.» Because they are based on discretionary considerations and, in some cases, merely recommendatory arguments, they are referred to as preference rules. A preference rule no longer provides clear criteria for application, but merely recommendations as to which decision is preferable depending on the situation. In simple cases, preference rules can also be used to derive clear guidelines for a decision. However, if several preference rules can be applied at the same time, it is quite possible that these criteria, which may conflict with each other, no longer allow for unambiguous decisions.
Intuitively, preference rules are a major advance in music theory because they can represent the ambiguity of musical structures much more appropriately than traditional rigid sets of rules.
Some examples of preference rules for group structures in GTTM:
The preference rules derive their criteria from, among other things, Gestalt theory, which plays an important role in cognitive psychology. The theory goes back to the works of psychologists Ernst Mach and Alexius Meinong, which were written at the end of the 19th century. However, the essay “Über «Gestaltqualitäten» by the Viennese Christian von Ehrenfels is considered to be the initial spark for its development. Von Ehrenfels refers explicitly to music. A melody, he says, is more than «the sum of its individual tones or intervals.» A melody has a certain form that is not destroyed even when transposed. Forms, on the other hand, are «mental images» and, as such, the creative work of the mind.
Gestalt theory later split into two directions: the Berlin School, influenced by the work of Max Wertheimer, Wolfgang Köhler, and Kurt Koffka, which also included Carl Stumpf’s assistant Erich von Hornbostel, and the Leipzig School around Wilhelm Wundt.
In 1912, Max Wertheimer described the so-called phi phenomenon. This is the apparent movement of objects that are in fact static. Wertheimer used the so-called Schumann tachistoscope to demonstrate this. Tachistoscopes are still used in psychological research today. They project an image for a freely selectable period of time, which can be as short as a few milliseconds.
Wertheimer projected two stripes alternately parallel and orthogonal:

The stripes, presented alternately horizontally and vertically,
suggested movement.
He found that the switching was perceived as movement. He transferred the results to music:
If specific optical phi phenomena occur, it should be mentioned that there are analogous problem areas in other sensory domains. For example, despite fundamental differences in the acoustic domain (…), “sound movement” as a characteristic, directed experience of a non-static nature shows some similarities.
This can be understood quite well when considering a simple melody:
![]()
Beginning of the «Ode to Joy» from Beethoven’s 9th Symphony
In fact, these are individual static sound events, but they appear to move as if a film were being played:

«Ode» as a sequence of film images
A single actor appears to move up and down, i.e., the individual, temporally distinct tones are identified as different states of a single object.
Gestalt theory principles form acoustic analogies to the corresponding findings in vision. As Köhler and Co. have established, the following points, for example, automatically break down into two groups in our perception because two of them are further apart:
°°° °°°°
The GTTM’s preference rules for group formation take such «natural» grouping tendencies into account.
However, GTTM does not only draw on the findings of Gestalt theorists. The book also sheds new light on the findings of Heinrich Schenker. The authors note that their prolongation analyses very often lead to a reduction that corresponds to Schenker’s fundamental theorem, but they also present examples that – according to common sense – do not originate from a fundamental theorem form. Chopin’s A major prelude from Op. 28, for example, does not fit the original theory. The prelude reduces to a chain of V-I steps that can only be squeezed into the I-V-I scheme of the original theory with a great deal of rhetoric (which the dogmatic Schenker would certainly have been capable of).
GTTM is the most subtle and consistent theory that has been presented to date with regard to musical-linguistic structures. There is certainly hope that it can be used as a basis for formulating a solid computer model for the automatic analysis of musical works (although one might ask what the point of such a model would be). However, there are also objections. One of the most significant could be described as a modern variant of the Galileo trap: the subtle tension/release and grouping structure focuses almost entirely on a melodic progression and cannot be transferred to polyphonic contexts. The authors themselves admit this: «In truly contrapuntal music, there is an important sense in which each individual line should have its own separate structural description. Although it is possible in principle to extend the theory to simultaneous structural descriptions, the formal complications would be so enormous that they would obscure the presentation of other, perhaps more fundamental aspects of musical structure.»
A polyphonic extension of GTTM has been promised, but to date it has only been touched upon fragmentarily, on the one hand by Fred Lehrdahl in his later book «Tonal Pitch Space», and on the other by his student David Temperley in the book «The Cognition of Basic Musical Structures». Fred Lehrdahl explains today that his interests have turned to other areas. However, he remains convinced that formulating a theory that encompasses polyphony—even if difficult—does not pose any fundamental problems. One may have doubts about this, because the interactions between the most diverse factors must become so subtle and complex that they would be almost impossible to grasp. Furthermore, it would first have to be proven that the human mind actually takes such complexities into account in its understanding of music.
Further objections arise from the heterogeneity of the framework. Two major problems arise here:
Problem number one: As Richard Cohn notes in a review of GTTM in the journal Theory Only 8 (1985), the GTTM rules have a circular character. Consider the above-mentioned preference rule for group structures and another for time span reduction, which is also found in GTTM. It reads as follows:
However, this means that the group structure is determined taking into account the time span reduction, and the time span reduction is in turn determined taking into account the group structure. This naturally smacks of a fatal circle.
The second problem lies in the nature of preference itself – the difficulty of reliably determining which rule carries greater weight when two preference rules conflict. For example, how can harmonic energy be weighed against gestalt perception? In many cases, the decision is intuitive, and it was precisely this – intuitive analysis – that GTTM originally sought to eliminate.
Ray Jackendoff made an attempt to get even closer to the process of hearing in a later article. In it, he asks to what extent theory actually tells us anything about what goes on in the human brain when we listen to and understand music. Jackendoff wants to show that music is understood in a way that parallels language: the music arriving in the brain is processed and interpreted («parsed,» as linguists say) by a modular «processor.»
A theory of musical perception, Jackendoff argues, should include the following:
According to Jackendoff, GTTM covers points 1 and 2 of this list.
In the article, Jackendoff puts forward the following thesis on point 3:
During real-time analysis of music, the processor in the brain considers several possible analyses in parallel until one of them emerges as the only possible one.
This thesis makes several assumptions. First, it transfers the model of the linguistic parser to the analysis and understanding of music. Jackendoff rules out two alternatives to the parallel determination of analyses. One alternative would be that only one (the most plausible) analysis is carried out until it loses plausibility to such an extent that it is discarded in favor of another. Jackendoff argues that this would require several analyses to be available at the outset, from which the parser could choose one. Furthermore, if the analysis failed, the parser would return to the beginning to start a new analysis. Since this happens in real time, the musical event would continue, and the parser in the brain would have to analyze what has already been analyzed again and also take the current musical event into account. This would very quickly lead to the parser becoming overloaded.
A second alternative would be for the parser to leave the analysis open until it reaches a point where the musical event allows for an unambiguous interpretation. Jackendoff calls this the serial-indeterministic model. His objections to this are that in order to apply the correct analysis at the moment of unambiguity, the parser must already have all analyses available. However, Jackendoff dismisses this as fortune-telling:
It is merely wishful thinking to assume that the correct analysis will somehow miraculously «appear» once enough evidence has been accumulated. If we want to explain the principles of correct analysis in the language of information processing, it is necessary for the parser to formulate several possible analyses from which the correct one can be chosen. If this is not the case, then we are essentially claiming that the mind does this in a magical way – that is, we are denying a rational explanation.
A second problem with this type of parser is that it makes the wrong kind of predictions about the nature of musical experience; it suggests that we (…) do not experience any metrical structure at all [as long as it is not unambiguous], because the parser does not choose one until then. This is obviously wrong: we clearly have a metrical intuition long before we have definitive proof of it.
The parallel analysis of the thesis favored by Jackendoff encounters difficulties that he himself apparently overlooks. Jackendoff assumes that there is always a finite, in principle known set of analyses to choose from. He compares this to the analysis of ambiguous words in sentences, such as «bank,» which can mean a sandbank in a river, a place to sit, or a financial institution. In this case, experiments have shown that the listener has all possible meanings in mind until one of them turns out to be the correct one.
Unfortunately, however, there are also sentence beginnings in language that allow for an indefinite and undefined number of analyses. We German speakers can sing a song about this thanks to the sentence bracket:
Yesterday, lost in thought, I (…) my bread …
So what happened? …ate it too early? …packed it away? …gave it to a child? …fed it to a horse? …shot it to the moon?
Even if we don’t know how the sentence will end and we have no way of knowing which analyses to choose from, we already interpret the parts that are available. Since the constructive rules in music are subject to a certain degree of arbitrariness, we are dealing with a similar situation: we may not have a clear and known set of analyses at our disposal. Such considerations can lead to a fundamental rejection of GTTM as a model of human cognition. However, the merits of the two authors remain undisputed when one considers the subtle way in which GTTM allows musical structures to be analyzed.
GTTM is, all things considered, the culmination of the musical-grammatical metaphor of music and, at the same time, its dissolution. The consistent application of grammatical framework theories ultimately leads to the transcendence of the language metaphor as such. The grammatical metaphor only comes into play in the form of jargon. When Lerdahl and Jackendoff describe music and musical structures informally, they do so primarily with terms borrowed from physics and spatial theory: melodies go up and down, there are centers of gravity and energies, and all kinds of forces are at play.