I rewrote my previous blog post twice

I love to write. Decades after I learned to read, it still amazes me that the written word can cross miles and millenia to transmit a writer’s thoughts to his or her readers. In high school, a series of excellent teachers taught me that writing is an expression of thinking, so that if you’re having trouble writing something, chances are you haven’t thought it through thoroughly enough.

I started this blog in 2013 in part to research and think through topics that I would later include in my first book. It was a helpful exercise and kept me motivated. It was also interesting to me at a meta level. To my surprise, when I eventually converted blog posts to book sections I often inverted them. My blog posts, which were discursive in nature, often ended with a twist. My book sections, which were expository, often turned that twist into an introductory hook.

Now that I’m working on a third book I’m again finding it useful to post as I carry out my research. It’s highly motivating to be thinking, as I read and take notes, “Is there anything here interesting enough to blog about?” As with my first book, I expect that many of these posts will make it into my book in some form. And the craft of writing has continued to surprise me.

My most recent post, “Coriander, Cilantro, Culantro“, which I published two weeks ago, is a case in point. I wasn’t thrilled with it at the time but was anxious to post something. It had been almost two weeks since my previous post, which wasn’t even related to my book research. However, afterwards I found that I couldn’t move on in my research, but instead kept coming back to the “Cilantro” post. It was like compulsively probing a sore tooth. The more I reread it, the more I was convinced that it was boring. I had presented the three herbs through the prism of my own life experience (who cares?), and only in the second half of the post had I explained why the three words were of linguistic interest. And the post didn’t explain why, seemingly out of the blue, I was writing about herbs.

So a week later I completely revised the post. This second version added the necessary context, and explicitly addressed the linguistic interest of this triplet of words, but did so clunkily, as a series of bullet points.

As a result I found myself stuck again, and finally, today, posted a revision that satisfies me. It’s the one you’ll see on this blog now. It puts the linguistic information up front, but with, I think, a good structure and flow. It adds my personal perspective, but does so lightly. And I worked in a twist at the end.

I’ve memorialized the first two versions of my “cilantro” post here and here. Have a look if you like, and let me know if you agree that the current one is really superior. Now I can finally move on. Unfortunately, the next word in Histories and Mysteries is…COCKROACH!!!

Coriander, cilantro, culantro

Slogging through the letter C in Spanish Word Histories and Mysteries and taking notes for my third book, I was interested to learn that a single Greek word, koriandron, gave rise to the English and Spanish words for the three different culinary flavorings pictured below: coriander seeds, which are sold both whole and ground, and two leafy herbs: cilantro and culantro. Cilantro is the herb that grows from coriander seeds, and culantro is a botanical cousin. Like cilantro, it is a member of the apiaceae family, which also includes carrots, celery, and parsley.

By the way, I use all three of these flavorings in my own cooking. I prefer to grind coriander seeds myself, as needed, for Indian recipes (Madhur Jaffrey is my guru). I use cilantro constantly. And recently I have been using culantro, which I buy at H-Mart, in Costa Rican recipes from The Blue Zones Kitchen. And now, back to linguistics…

Greek words are most common in technical domains, such as science (e.g. cometa) and medicine (e.g. epilepsia), as well as religion (e.g. biblia). In the kitchen, French and Italian borrowings predominate for obvious reasons. However, according to Histories and Mysteries, Greek is the source of many names of herbs. This makes sense because herbs straddle the boundary between food and medicine

Latin borrowed Greek koriandron as coriandrum. In French coriandrum developed into coriandre, which is the source of English coriander. In Spanish coriandrum developed into culantro. This word was first attested in 1385; cilantro (first attested in 1680) developed as a variant pronunciation but eclipsed culantro in frequency in the 20th century (see Google ngram analysis below). Perhaps this is when culantro became associated with the less common herb. At any rate, English borrowed both cilantro and culantro from Spanish to complete its herbal trio.

According to the authoritative Spanish etymologist Joan Corominas, the transitions from coriandrum to culantro, and from culantro to cilantro, were both noteworthy.

  • In the first phase, the change of the first /r/ in coriandrum to an /l/ is an example of ‘dissimilation’: a sporadic process whereby nearby sounds became more different from each other in order to make a word easier to pronounce. A similar change occurred in the evolution of Vulgar Latin robore into Spanish roble ‘oak’, and of Latin peregrinus into French lerin. (The other changes, such as final -um becoming -o, happened across the board as Spanish evolved.)
  • Regarding the second phase, Corominas writes that the vowel change from the /u/ of culantro to the /i/ of cilantro was unique. Since /i/ and /e/ are similar vowels, he suggests that the cilantro variant might have been influenced by another botanical name: celidonia (celandine in English, botanically Chelidonium majus), a plant whose leaves are toxic if ingested, and whose sap has been used as a folk remedy for warts.

You say cilantro, I say culantro — let’s cook!

Why is rr different from ch and ll?

Yesterday I had a linguistic panic attack at my gym. I always speak Spanish with the staff there, who are are uniformly Latina, and they know that I’m interested in the language. Yesterday one of them asked me why Spanish has more letters than English. This was a mistake on her part, because I immediately launched into an explanation of how the Real Academia Española (more accurately, a majority of the world’s Spanish language academies) eliminated ch and ll from the alphabet in 1994.

Afterwards the question popped into my head: what about rr, the Spanish double r that denotes a trill? Along with ch and ll, rr is the language’s third purely consonantal digraph, meaning a sequence of two consonant letters that together represent a single consonant sound. Was rr also affected in 1994? (I was 99% sure that it wasn’t.) If not, why not? And why hadn’t I thought about this before? (That was the panic part.) I bailed on my workout early.

Once I got home and consulted my bookshelf and the Internet, I quickly reassured myself that rr had never been part of the Spanish alphabet, and so had not merited the same scrutiny as ch and ll. Any easy proof of rr‘s inferior status is that words containing an rr have always been alphabetized as they would be in any other language. For example, in the 2006 compact Berlitz dictionary that I still keep on my shelf (but rarely use these days), corrupción precedes corsetería and perro precedes persecución. You will find the same order in an older dictionary. If rr were part of the alphabet, between r and s, then all cor- words would precede all corr- words, and all per- words would precede all perr- words.

That is exactly how ch and ll were treated before 1994. In a modern dictionary pecho ‘chest’ precedes pecoso ‘freckled’ and callar ‘to be quiet’ precedes calor ‘heat’. However, in pre-1994 dictionaries, pecoso preceded pecho and calor preceded calle. Because the digraphs ch and ll were then considered full-fledged elements of the Spanish alphabet, all pec– words preceded all pech– words, and all cal- words precede all call- words.

In fact, this uniquely Spanish version of alphabetical order motivated the 1994 alphabet reform. As I wrote in my first book,

The change in the alphabet...simplified language processing software. Before 1994, software routines that relied on alphabetical order needed to make exceptions for Spanish, since all Spanish words with c were alphabetized before those with ch, and likewise for l and ll. For example, culebra ‘snake’ was considered alphabetically prior to chico ‘boy,’ and luchar ‘to fight’ prior to llano ‘flat, plain’—a bizarre phenomenon from a non-Spanish perspective. With ch and ll eliminated as separate letters, Spanish no longer required special computational handling. [From Question #45, "What happened to ch and ll?"]

The question remains: why was rr never treated like a letter in the first place? I don’t know (maybe some reader does), but I have a theory. In written Spanish, words can begin with any consonant, and also with ch or ll, but not with rr. Because a single r at the beginning of a word is pronounced as a trill, there would be no difference between, say, rojo ‘red’ and a hypothetical rrojo. In this sense the ch and ll digraphs were more salient than rr. Surely this is why they were added to the alphabet.

Catching some z’s

By now I’m well into the “c” section of Spanish Word Histories and Mysteries. This is decent progress, since in my previous post inspired by this book I was only in the “b’s”. However, my perusal ground to a halt yesterday when I read about the English word cedilla, which is a straight borrowing of the Spanish word cedilla. The cedilla is the little hook that some languages use under a letter, most often the letter c, to represent certain pronunciations. In French the cedilla softens a c to a /s/ before a or o, as in français ‘French’ or garçon ‘waiter’. A plain (no-cedilla) French c is pronounced /k/ before these vowels, as in carte ‘menu’ or couper ‘to cut’.

The cedilla is not used in Modern Spanish, but in Old Spanish ç was pronounced /ts/, as in the middle of pizza. As in French, ç was only used before a or o, for example in Old Spanish caçar ‘to hunt’ (now cazar) and alço ‘I raise’ (now alzo). As Old Spanish became Modern, the sound /ts/ disappeared, and words previously written with a ç switched to a z, as shown above, that was (and is) pronounced either s or th, depending on one’s dialect.

Histories and Mysteries explains that the cedilla under ç originated as a miniature letter z. This is apparent if you look at a large ç (below). In Spanish, ceda is a variant spelling of zeta, the name of the letter z, so cedilla, with the diminutive -illa suffix, simply means ‘little z‘.

I never knew this! and it makes me enormously happy because the tilde (~) over the beloved Spanish ñ originated as a small letter n. What an awesome parallel! But it also made me realize, for the first time, the irony of zeta, the Spanish word for the letter z.

Why ironic? Because one of the rules of Spanish spelling is that z is only used before a, o, and u, and is replaced by c before e or i. You can see this in the Spanish spelling of cebra and cero (below), and in spelling alternations like empezar ‘to begin’ and empiezo ‘I begin’, versus ¡Empiece! ‘Begin!’ and empecé ‘I began’. Most words that violate this rule are recent borrowings like zepelín ‘zeppelin’, zenit ‘zenith’, zig-zag, zinc (the metal) and zigoto ‘zygote’. Ironically, the word zeta is another exception.

As far as I know, this rule lacks a logical explanation, since the pronunciation of the letter z isn’t affected by the vowel that follows it. The c/z alternation is thus different from the language’s other spelling alternations, which are well motivated. For example, adding a u to a g in verb forms like pagué ‘I paid’ preserves the hard /g/ sound of the verb pagar. For this reason I’m afraid I used to tell my students that it was “a stupid rule of Spanish” to always use c, not z, before e and i. Honestly, I’ve looked for explanations, but have never found one that satisfies me.

So the “irony of zeta“, alluded to above, is simply that its spelling violates the Spanish spelling rule for the letter z. To put it another way, you can’t state the rule “Don’t use the letter z before e or i” without breaking it: No se usa la letra zeta antes de e o i.” What an oxymoron! I don’t know if cedilla will be one of my 100 words, but I must mention this irony.

By the way, although zepelín and zig-zag seem to be maintaining their alien z spellings, the hispanized spellings cinc and cigoto have threatened or surpassed the originals, as shown in the Google ngram results in the slideshow below. This shows that the c/z spelling rule, motivated or not, maintains its force in modern Spanish.

Much ‘-ado’ about nothing

Yesterday I made good headway in Spanish Word Histories and Mysteries: thirteen words, many of which will be useful grist for the mill of my next book. I’ve titled this post “Much ‑ado about nothing” because the -ado or -ada endings of three of these words — armada, bastinado (from Spanish bastonada), and avocado — make them look like Spanish past participles, but two of them are faking it.

A grammatical aside: a “past participle” is a verb form like cerrado ‘closed’ in the sentence He cerrado la puerta ‘I have closed the door’. A participle can also be used as an adjective, as in La puerta está cerrada ‘The door is closed”, in which case it must agree with the noun it modifies (puerta is feminine and singular). Compare libros cerrados ‘closed books’, where noun and adjective are both masculine and plural.

  • Armada did, in fact, originate as a past participle, although it is obviously a noun. It means ‘navy’, and in English it normally refers to the famous Spanish Armada, the enormous fleet of Spanish ships that in 1588 attempted to spearhead an invasion of England, but lost decisively. (An English “counter-armada” against Spain the next year, led by Sir Francis Drake, failed as well.) Armada is the feminine form of the past participle of the verb -armar ‘to arm’, and therefore means ‘armed’. In Spanish one often sees it in the phrase fuerza armada ‘armed force’ or its plural, fuerzas armadas ‘armed forces’.

    I was curious to know why the noun armada is feminine. Is this because of its frequent usage as a feminine adjective, as shown above? The answer turns out to be “sort of.” The direct antecedent of armada was the feminine plural form of the LATIN past participle armatus, found commonly in naves armatae ‘armed ships’. So: feminine by association, yes — but in Latin, not Spanish.

  • English bastinado has a few related meanings. It can refer to beating someone with a stick, a single blow with a stick, or a stick used for beating people. It is based on the Spanish word bastonada, which appears to be another feminine-past-participle-turned-noun, this one based on a verb bastonar. However, this verb doesn’t exist, and never did. Rather, bastonada was formed from the noun bastón ‘stick’ (related to the French word baton) plus the suffix -ada, which Spanish uses to derive nouns that describe physical violence. A cabezada (from cabeza ‘head’) is a header in soccer or a headbutt in a fight. A cuchillada (from cuchillo ‘knife’) is the cut that a knife makes: a stab or a slash. A puñada (from puño ‘fist’) is a punch. And so on.

    Someone who doesn’t know Spanish might be surprised to learn that it has such a specific suffix. For me, the surprise was learning that it has two. I already knew about its -azo suffix, which has the same meaning but is more productive, in the linguistic sense of being freely applied to almost any noun of the speaker’s choice. The Real Academia (previous link) comments that because of this productivity, los diccionarios no puedan recoger todas las voces admisibles así formadas — that is, dictionaries cannot possibly include all the words of this type. For example, if you hit someone with an elephant, this would be an elefantazo.

    In fact, three out of the four -ada ‘blow’ words above have more common -azo cousins: bastonazo, cabezazo, and puñetazo. Only cuchillada beats its cuchillazo alternative, although when we talk specifically about navajas (pocket knives) navajazo is more common than navajada.

  • Finally, the English word avocado looks like another noun based on a past participle. However, it comes instead from the Spanish word aguacate, borrowed from Nahuatl ahuacatl. The -tl consonant cluster, seen also at the end of the language name itself, was difficult for Spanish speakers to pronounce, and changing it to -te was a common solution. Thus tomate, the Spanish word for tomato, comes from Nahuatl tomatl.

    Why did English adapt aguacate as avocado? According to Histories and Mysteries, avocado was a variant form used by Spanish speakers before aguacate emerged as the standard. Elsewhere this variant form is attributed to ‘folk etymology’, i.e. the word form being assimilated to a similar pre-existing word, in this case abogado ‘lawyer’. In French, ‘avocado’ and ‘lawyer’ ended up identically, as avocat.

    All this was familiar to me before I read the avocado entry in Histories and Mysteries. But I was happy to learn that guacamole comes from the Nahuatl word ahuacamolli, literally ‘avocado sauce’. The second part of the word, molli ‘sauce’, is the source of the culinary term mole (the Mexican sauce). I have to confess that for decades I pronounced this word moLÉ, to rhyme with the exclamation ¡Olé!. This certainly captures how I feel when I eat chicken with mole sauce.

I didn’t forget the paragraph

My (adult) children still remember that when they they were schoolchildren, if they left the house in the morning but ran back a minute later crying “I forgot my lunch!” I would always say, “You didn’t FORGET your lunch, you REMEMBERED your lunch”. This undoubtedly got old after a while.

Anyway, today I realized that I forgot to include a vital paragraph in yesterday’s post, about the ways that words in Spanish changed once they entered the lexicon. I have now added the paragraph; you can find it easily if you search down for the phrase “Once each word.”

So…I didn’t forget the paragraph, I remembered it. Enjoy!

Crystal Clear

After more months than I ever would have imagined, I have finally finished taking notes on David Crystal’s The Story of English in 100 Words. As I had expected, this book was an invaluable first step in my research for the analogous book about Spanish that I plan to write. Now I will move on to the American Heritage Dictionaries’ Spanish Word Histories and Mysteries, into which I dipped my toe (and my pen) some days ago, when I had temporarily misplaced The Story of English.

When I sat down to scope out my new book I wrote out a list of several dozen words that I wanted to include. This list contained specific words like yugo and canoa as well as placeholders like “a word from Galician/Portuguese” or “a name of a Latin American country.” As I worked through Crystal’s book this list mushroomed to 172 words. This isn’t a problem: although each entry in the book will focus on one word, it will also discuss other words that are relevant in some way. In fact, I expect that the list will continue to grow as I do more research. The real challenge will be condensing the list into the ideal 100 words, each with its supporting cast.

My main goal in choosing those words is still to capture the entire history of the Spanish language. This organization will lend itself naturally to covering the full range of languages that have contributed to Spanish, from its Latin core to modern borrowings from Asian languages and technology.

While a canonical borrowing adds a new word for a new concept, such as canoa, the Arawak word for a kind of boat that the Spaniards encountered in the New World, my list includes borrowings that are more complex, such as:

  • calques: words that translate a foreign term part-by-part into Spanish (e.g. tranvía, modeled on English tramway)
  • words that passed through multiple languages en route to Spanish, such as adobe (see previous post);
  • doublets: borrowings that co-exist with older words from the same source, e.g. limpio ‘clean’ and límpido ‘clear, limpid’, which both came from the Latin word limpidus. Limpio is an older word that came from Vulgar Latin, whereas límpido is a 19th-century learned borrowing from the same Latin root.
  • borrowings that were adapted to Spanish pronunciation (biftek), that are pronounced according to the rules of Spanish spelling (club, pronounced with a Spanish /u/), or that expanded the horizons of Spanish pronunciation (the tl- in chipotle, or the Mexican Spanish tlapalería ‘hardware store’).

I’d like to address, if possible, the general question of why Spanish, like English and in contrast to French, has been so accepting of borrowings. But I wonder: how can I even pose this question without anthromorphizing “Spanish”?

Many Spanish words are neither from Latin nor borrowed from another source. My list includes examples from a myriad of other sources and processes, such as:

  • names of people (biro), places (jerez), or commercial brands (celo), both real and fictional (quijotesco)
  • onomotapoeia (chirriar)
  • truncation (moto)
  • texting (jajaja)
  • compounding (lavaplatos)
  • affixation (desencadenar)
  • back-formation (debate, from debatir)
  • phrases (usted, pordiosero)
  • blends (agridulce)
  • acronyms (sida)

Once each word entered the Spanish lexicon, it was not frozen in form and meaning, but rather was vulnerable to the normal currents of language evolution. My list includes words that changed their form. For example, harina ‘flour’, which began with an f- in Latin (farina), dropped the initial /h/ sound between Old (medieval) Spanish and Modern Spanish, although the letter /h/ was maintained in spelling. The same happened to other words of this type such as hierro ‘iron’ (from Latin ferrus) and higo ‘fig’, from Latin ficus. I also include words that changed their meaning in a variety of ways. Some:

  • became more negative or positive, such as siniestra ‘sinister’ and choza ‘house’, which originally meant ‘left’ (the direction) and ‘hut’;
  • became more general or narrow in meaning, such as pariente ‘relative’ and infante ‘prince’, which originally meant ‘parent’ and ‘child’.
  • acquired new meanings as the culture changed, such as pavo, whose meaning changed from ‘peacock’ to ‘turkey’ when the latter was encountered in the New World;
  • absorbed meanings from other languages, such as ratón ‘mouse’, which absorbed the meaning of the English computer ‘mouse’;

Most words on the list are nouns and verbs, but there are also adjectives, prepositions, pronouns, adverbs, and exclamations. The list includes examples of dialectal variation in meaning (tortilla) and pronunciation (cerveza).

Spanish spelling is fairly straightforward, but I have included words that illustrate

  • the use of accent marks
  • the ñ
  • spelling changes in conjugated forms (busqué ‘I looked for’, from buscar ‘to look for’)
  • spelling adaptations for casual language (pa’ for para)
  • texting abbreviations (q or k for que)
  • gender-neutral language (amig@s)
  • words whose place in alphabetical order changed when the Real Academia eliminated ll and ch as digraphs in 1994. (The change was promulgated in its 2010 revision of its Ortografía, or spelling manual.)

As I converge on a definitive list I hope to represent most of the two dozen semantic categories in the anthropology-based World Loan Database (WOLD), such as kinship terms, food and drink, motion, and spatial relations. Based on Crystal’s book I’ve added other categories such as swear words, politeness expressions, words for unknown people or things (fulano, chisme), and titles (señor).

Finally, my ideal final set of words will illustrate fundamental properties of Spanish grammar and its evolution, such as verb forms based on haber, verbs with irregular conjugations and nouns with irregular gender, reflexive verbs, the various origins and pronunciations of ll, and pluralia tantum (like tijeras).

It will be interesting to see how my list continues to grow and change as I work through additional sources.

A detour into a new resource

It’s been so long since I’ve worked on my new book (my current excuse is excessive travel) that when I finally sat down yesterday to get back to work, I couldn’t find my copy of David Crystal’s The Story of English of 100 Words, which I’ve been taking notes on for ages. I found it later in the day, but in the meantime spent my allotted work time taking notes on the first several pages of Spanish Word Histories and Mysteries, which I have in mind as my next resource. While this book is really about English, it also serves as an approachable and well-written introduction to Spanish word origins. It discusses 150 Spanish words, from A to Z, that have made their way into English.

Yesterday I only worked through the book’s Introduction, and the beginning of the section devoted to the letter “A”, but this short detour immediately bore fruit. For one thing, what I read broadened my thinking about the origin of Spanish words for indigenous American flora, fauna, and so on. I knew that Spanish borrowed many words from the languages of South American, Central America, and Mexico, but I hadn’t given much thought to the languages spoken in what is today the United States. Yesterday I learned that the English word abalone comes from the Spanish word abulón, itself derived from aulun, the word for abalone in Rumsen, a now-extinct Ohlone language spoken in today’s Monterey County. Abalone was an important part of the Rumsen diet.

Spanish Word Histories and Mysteries dates this borrowing to the middle of the 1700s, when Spain became an active presence in the San Francisco Bay Area, building Franciscan missions and civilian settlements. I already knew that Spain successively borrowed words from each region of the Americas as they conquered it: first the Caribbean, then Mexico and the Pacific Coast of South America, and finally the South American interior. (See Table 4.1 from my first book, copied below). Learning about abalone/abulón/aulun reminded me to include Ohlone and other languages spoken in today’s United States in the latter group.

Not all Spanish words for American flora and fauna (etc.) come from indigenous American languages. The Introduction to Spanish Word Histories and Mysteries uses a fun pair of words to illustrate this point. While chocolate comes from Nahuatl (see again Table 4.1), vainilla is pure Spanish, the diminutive of vaina ‘sheath’. Spaniards used this word to name this unfamiliar plant because of the shape of the vanilla pod or vanilla bean.

Many words in the book’s “A” section are nouns borrowed from Arabic with the article al ‘the’ prefixed to the noun root. Adobe, from Arabic at-tûba ‘brick’ (here the l of al assimilated to the following t), is an example of an Arabic word that itself was borrowed from another language. Tûba comes from Coptic, a late form of the Egyptian language, which is only distantly related to Arabic. (It’s an Afro-Asiatic language, but not in the Semitic branch of the family). I’m familiar with other examples of Arabic borrowings that come from other languages, such as alambique ‘still for making alcohol’, from Greek, and aduana ‘customs officer’, from Persian.

Finally, spending time with yesterday’s words reminded me that vocabulary is as interesting as the world it describes. Despite living in the Bay Area for several years, I hadn’t heard of the Ohlone people. Nor did I know that the orange color of cheddar cheese normally comes from annatto, a pigment derived from achiote (originally a Nahuatl word). On a personal level, because I am Jewish it was satisfying to learn that adobe has Egyptian roots; every year at the Passover seder we remember that the Egyptians forced their Jewish slaves to build cities for Pharaoh “with harsh labor at mortar and bricks”. (However, the Hebrew word for brick used in the Book of Exodus (le’vey’nah) doesn’t appear to be related to tûba.)

Reading about the word albatross, from Spanish alcatraz, inspired me to learn about these amazing birds, long-lived and monogamous, who normally live in the air, where they can soar for hours on end, and on the water, where they find their food. They only land on the ground to mate, which happens annually or less often — every other year for the Great (or “wandering”) albatross. It was also interesting to learn that the negative meanings of albatross (the bird) and Alcatraz (the prison) are confined to English. (The former is attributable to Coleridge’s Rime of the Ancient Mariner.) For this reason English albatross is more interesting than Spanish alcatraz, which is unlikely to make the cut for my book.

I will continue with Crystal’s book, now that I’ve relocated it, but will look forward to returning to Mysteries and Histories next — hopefully, not too long from now.

Spelling pronunciation in Spanish

I’ve given up on often.

When I was growing up I was taught that the t in often is silent, and that pronouncing the word as ‘off-ten’ was a mark of a poor education. Merriam-Webster relates the silent t in often to those in soften, hasten, and fasten. Apparently the t in often was originally pronounced, but dropped out of vogue in the 1500s, so that Queen Elizabeth, for example, avoided it. M-W writes that “phonetically spelled lists made in the 17th century indicate that ‘the pronunciation without [t] seems to have been avoided in careful speech'”, and that “three hundred years later, the first edition of the Oxford English Dictionary” said that the off-ten pronunciation was “not recognized in dictionaries.”

However, the 1934 unabridged Webster’s Second reported that the off-ten pronuncation, “until recently generally considered as more or less illiterate, is not uncommon among the educated in some sections,” and the current M-W dictionary “reports both pronunciations as equally accepted.”

Indeed, in my lifetime I have observed more and more people using the off-ten pronunciation regardless of education. It is certainly much more common today than the pronunciation with a silent t. I still grit my teeth when I hear off-ten, but smother the impulse to think less of the speaker.

What does this have to do with Spanish? In an effort to get back to work on my new book, after taking some time off due to a death in the family, I have returned to my previous task of gleaning inspiration from David Crystal’s The Story of English in 100 Words, and in fact have hit the halfway point (50 words). Crystal’s word #40 is debt, which of course has a silent b. Crystal explains that debt, which was borrowed from French dette, was originally spelled det or dett in English. In the 16th century aspirational English speakers added a b in imitation of the Latin root debitum, with support from the related English word debit. Cyrstal provides several other examples of this hyper-correct “spelling reform”, such as the b in doubt and subtle and the p in receipt.

Crystal also points out that although the added letters in debt, doubt, subtle, and receipt are still silent, English speakers came to pronounce some letters added in the 1600s, such as the s in baptism and the l in fault. Although Crystal didn’t use the term, linguists describe pronunciations like these as “spelling pronunciation”, i.e. pronouncing words the way they are spelled. Off-ten is another example of spelling pronunciation.

Crystal’s examples of spelling pronunciation caught my eye because I’m hyper-aware of how people pronounce often. However, I assumed that spelling pronunciation would be irrelevant to Spanish because Spanish spelling is phonetic: that is, the pronunciation of any written Spanish word is obvious and unambiguous.

Nevertheless I googled “Spelling pronunciation in Spanish” just to be thorough. To my surprise I ran into several examples, all of which involve words borrowed from other languages. Usually Spanish changes foreign spelling to conform to Spanish principles, using different letters, as in bistec ‘beefsteak’, and/or adding accent marks as in mánager and esplín ‘spleen’. But some borrowings go the opposite direction, with Spanish speakers keeping the original spelling but pronouncing the foreign word according to Spanish rules. Some examples are:

  • clóset, with the English /z/ sound changed to /s/;
  • Mach (the scientific term, borrowed from German), with the final consonant pronounced with the ch of chorizo instead of the /x/ of ajo;
  • folclor, from English folklore, with the l in folk pronounced;
  • élite, from French élite, pronounced with three syllables (é-li-te), and with the French acute accent on the first /e/ interpreted as an indication of penultimate stress [I love this!];
  • iceberg, pronounced with three syllables (i-ce-berg) in Spain.

By the way, “spelling pronunciation” has always been one of my favorite bits of linguistic jargon because its meaning isn’t obvious from its two component words. (One might think it refers to phonetic transcription, i.e. a way of spelling words the way they are pronounced.) Perhaps the semantic relation between the two words is unusual.

¡Yamnaya!

After taking some weeks off for summertime fun with my family I am now back to research for my third book. I’ve continued to work my way through David Crystal’s The Story of English in 100 Words — still inspirational and intimidating. But I’ve also taken a detour to devour Laura Spinney’s book Proto: How One Ancient Language Went Global, about the pre-history of the Indo-European languages. I will definitely want to touch on this topic in my book, since it’s part of the history of Spanish.

For those of you who may be unaware, the Indo-European language family includes thousands of languages spoken today: Romance languages like Spanish, Germanic languages like English, Slavic languages like Russian, Baltic languages like Lithuanian, and Celtic languages like Welsh; Greek, Albanian, and Armenian; Hindi and other languages of northern India; and languages of Central and Western Asia including Pashtun (Afghanistan) and Persian (Iran). It also includes languages that are no longer spoken and (unlike Latin) have no spoken descendents. These include Oscan and Umbrian ( “Italic” languages related to Latin), languages of ancient Turkey including Hittite and Phyrgian (“Anatolian” languages), and Tocharian A and B, once spoken in the Tarim Basin of northwestern China.

I learned about the Indo-European language family as an undergraduate and graduate linguistics student. However, my knowledge of the family’s origins was vague: I knew only that it arose somewhere north of the Black Sea. (Or was it the Caspian Sea???) Although I was aware that in the years — decades, really 😉 — since I completed my PhD, there had been substantial new research on this topic, I hadn’t followed the field.

Proto has brought me up to date — and can I say “Wow?” Spinney’s book combines recent research from linguistics, archaeology, and genetics (DNA analysis) to illuminate what is known today about the origin of the Indo-European languages. Essentially, Proto-Indo-European — the common ancestor of the language family — spread east and west through the Eurasian steppe (grasslands), conveyed by cattle herders equipped with horses and wagons. The steppe was a natural environment in which such people could thrive and spread. Researchers call these steppe herders the Yamnaya, which means ‘pit grave’ in Russian, because of their burial practices.

The illustration below, from pp. 60-61 of Spinney’s book, shows the progression of Indo-European (black arrows) through the steppe (shaded area) and beyond. My vague memory wasn’t that far off, since the arrows show a specific point of origin in the steppe north of the Black Sea.

Genetic analysis of DNA from Yamnaya remains suggests an exciting possibility:

“The earliest Yamnaya males … carried a very narrow cluster of Y chromosomes. Later on, after the population had grown, other Y chromosomes entered the mix, but the first of their kind may have been closely related to each other on their father’s side. One way to interpret the evidence is to think of the Yamnaya as a single clan or brotherhood who distinguished themselves by their burial rite. They may have left a larger group, or been expelled from it, and having moved out of their ancestral valley became increasingly nomadic unril they vanished into the grasslands for good. If that is who they were it prompts an extraordinary reflection: fewer than a hundred people may have spoken the dialect that gave rise to all extant Indo-European languages.” (p. 69)

While Proto doesn’t have anything to say about Spanish specifically, it does include interesting speculation about the connection among Italic, Germanic, and Celtic languages:

Germanic, Celtic and Italic are related by common descent. This is evident from their grammar, their pronunciation and their core vocabulary (English – – father mother brother; Old Irish athir máthir bráthir ; Latin pater mter frter). But the relationships between them aren’t equal. Celtic and Italic are generally considered to be closer to each other than either is to Germanic, like twins with a third sibling.…Some linguists suspect that Italic and Celtic arose as a single, possibly short-lived language, Italo-Celtic, while Germanic arose separately.

As I had hoped, Proto provided me with the information I need to write about Indo-European origins. On other counts, as a linguist I especially enjoyed learning about Anatolian and Tocharian, two branches of the Indo-European family that had always flown under my cognitive radar. I think that many people interested in languages or history would find the book to be an informative and accessible read. It provides a dizzying overview of the early history of humanity around the Black Sea: a region that, like the Fertile Crescent, would be a springboard for widespread advances in civilization for millenia to come. I recommend it highly.