Tag Archives: Martin Haspelmath

Crystal Clear

After more months than I ever would have imagined, I have finally finished taking notes on David Crystal’s The Story of English in 100 Words. As I had expected, this book was an invaluable first step in my research for the analogous book about Spanish that I plan to write. Now I will move on to the American Heritage Dictionaries’ Spanish Word Histories and Mysteries, into which I dipped my toe (and my pen) some days ago, when I had temporarily misplaced The Story of English.

When I sat down to scope out my new book I wrote out a list of several dozen words that I wanted to include. This list contained specific words like yugo and canoa as well as placeholders like “a word from Galician/Portuguese” or “a name of a Latin American country.” As I worked through Crystal’s book this list mushroomed to 172 words. This isn’t a problem: although each entry in the book will focus on one word, it will also discuss other words that are relevant in some way. In fact, I expect that the list will continue to grow as I do more research. The real challenge will be condensing the list into the ideal 100 words, each with its supporting cast.

My main goal in choosing those words is still to capture the entire history of the Spanish language. This organization will lend itself naturally to covering the full range of languages that have contributed to Spanish, from its Latin core to modern borrowings from Asian languages and technology.

While a canonical borrowing adds a new word for a new concept, such as canoa, the Arawak word for a kind of boat that the Spaniards encountered in the New World, my list includes borrowings that are more complex, such as:

  • calques: words that translate a foreign term part-by-part into Spanish (e.g. tranvía, modeled on English tramway)
  • words that passed through multiple languages en route to Spanish, such as adobe (see previous post);
  • doublets: borrowings that co-exist with older words from the same source, e.g. limpio ‘clean’ and límpido ‘clear, limpid’, which both came from the Latin word limpidus. Limpio is an older word that came from Vulgar Latin, whereas límpido is a 19th-century learned borrowing from the same Latin root.
  • borrowings that were adapted to Spanish pronunciation (biftek), that are pronounced according to the rules of Spanish spelling (club, pronounced with a Spanish /u/), or that expanded the horizons of Spanish pronunciation (the tl- in chipotle, or the Mexican Spanish tlapalería ‘hardware store’).

I’d like to address, if possible, the general question of why Spanish, like English and in contrast to French, has been so accepting of borrowings. But I wonder: how can I even pose this question without anthromorphizing “Spanish”?

Many Spanish words are neither from Latin nor borrowed from another source. My list includes examples from a myriad of other sources and processes, such as:

  • names of people (biro), places (jerez), or commercial brands (celo), both real and fictional (quijotesco)
  • onomotapoeia (chirriar)
  • truncation (moto)
  • texting (jajaja)
  • compounding (lavaplatos)
  • affixation (desencadenar)
  • back-formation (debate, from debatir)
  • phrases (usted, pordiosero)
  • blends (agridulce)
  • acronyms (sida)

Once each word entered the Spanish lexicon, it was not frozen in form and meaning, but rather was vulnerable to the normal currents of language evolution. My list includes words that changed their form. For example, harina ‘flour’, which began with an f- in Latin (farina), dropped the initial /h/ sound between Old (medieval) Spanish and Modern Spanish, although the letter /h/ was maintained in spelling. The same happened to other words of this type such as hierro ‘iron’ (from Latin ferrus) and higo ‘fig’, from Latin ficus. I also include words that changed their meaning in a variety of ways. Some:

  • became more negative or positive, such as siniestra ‘sinister’ and choza ‘house’, which originally meant ‘left’ (the direction) and ‘hut’;
  • became more general or narrow in meaning, such as pariente ‘relative’ and infante ‘prince’, which originally meant ‘parent’ and ‘child’.
  • acquired new meanings as the culture changed, such as pavo, whose meaning changed from ‘peacock’ to ‘turkey’ when the latter was encountered in the New World;
  • absorbed meanings from other languages, such as ratón ‘mouse’, which absorbed the meaning of the English computer ‘mouse’;

Most words on the list are nouns and verbs, but there are also adjectives, prepositions, pronouns, adverbs, and exclamations. The list includes examples of dialectal variation in meaning (tortilla) and pronunciation (cerveza).

Spanish spelling is fairly straightforward, but I have included words that illustrate

  • the use of accent marks
  • the ñ
  • spelling changes in conjugated forms (busqué ‘I looked for’, from buscar ‘to look for’)
  • spelling adaptations for casual language (pa’ for para)
  • texting abbreviations (q or k for que)
  • gender-neutral language (amig@s)
  • words whose place in alphabetical order changed when the Real Academia eliminated ll and ch as digraphs in 1994. (The change was promulgated in its 2010 revision of its Ortografía, or spelling manual.)

As I converge on a definitive list I hope to represent most of the two dozen semantic categories in the anthropology-based World Loan Database (WOLD), such as kinship terms, food and drink, motion, and spatial relations. Based on Crystal’s book I’ve added other categories such as swear words, politeness expressions, words for unknown people or things (fulano, chisme), and titles (señor).

Finally, my ideal final set of words will illustrate fundamental properties of Spanish grammar and its evolution, such as verb forms based on haber, verbs with irregular conjugations and nouns with irregular gender, reflexive verbs, the various origins and pronunciations of ll, and pluralia tantum (like tijeras).

It will be interesting to see how my list continues to grow and change as I work through additional sources.

Something borrowed, something blue

For the last few years I’ve had a research project about Spanish word origins on the back burner. This summer I’ve resurrected the project, and it is simmering nicely: I have now finished the first major stage.

The focus of the project is Spanish borrowings, or loanwords: words in Spanish that originated in other languages. The project applies to Spanish the methodology from Martin Haspelmath and Uri Tadmor’s World Loanword Database (WOLD) project. Beginning in 2004, Haspelmath and Tadmor organized a team of linguists to collect data on loanwords in forty-one languages around the world. In 2009 they published their results in a book, Loanwords in the World’s Languages: A Comparative Handbook (De Gruyter), and the contributing linguists shared their data on the WOLD website.

My goals in this project are:

  1. To compare Spanish to the forty-one languages in the WOLD project, in terms of (i) its percentage of loanwords, and (ii) these words’ characteristics, such as their part of speech.
  2. To quantify the relative contributions of different source languages to Spanish vocabulary. I already did this for my first book, using a random sampling of five hundred words from a standard Spanish etymological dictionary. But that sample may have skewed toward more recherché vocabulary.
  3. To address various issues involved in etymological research, in Spanish and in general.

More about the WOLD project

In order to obtain comparable results across the WOLD languages, all participating linguists started with the same list of 1460 core meanings: ‘house,’ ‘mother,’ ‘go,’ and so on. Each linguist identified ‘their’ language’s words for these meanings, then traced the origins of those words using a standardized set of guidelines. I have now completed the first of these two steps for Spanish. It raised all sorts of interesting issues, which I will discuss in my next blog post.

One goal of the WOLD project was to compare the frequency of borrowing in different languages. In other words, of the core meanings, how many were expressed in each language by loanwords? As shown in the table below, borrowing rates ranged from 1.2% for Mandarin Chinese to 62.7% for Selice Romani. Yaron Matras’s review of the WOLD Handbook in the journal Language points out that these two languages are spoken in diametrically different environments. Speakers of Mandarin “show little or no bilingualism”; the language has “a status as a majority language, a powerful standard, and a sociopolitically dominant population.” In contrast, Selice Romani is associated with “universal multilingualism, a minority language status, the absence of a written standard, and sociopolitical marginalization.”

Romanian, the only Romance language in the project, fell into the “high borrowers” category (25.9% to 45.6%), as did English. My previous research (see above) placed Spanish in the “very high borrowers” category, with roughly one-third “native” vocabulary (from Vulgar Latin), one-third later borrowings from Latin, and one-third words from other languages. It will be interested to see whether this holds up for a WOLD-based lexicon.

Borrowing typeLanguages (in increasing order of % loanwords)
“Low borrowers”
(1.2 – 9.7%)
Mandarin Chinese, Old High German, Manange, Ket
“Average borrowers”
(10.7 – 22.4%)
Otomi, Seychelles Creole, Gawwada, Hug, Oroqen, Hawaiian, Kali’na, Iraqw, Q’eqchi’, Wichí, Zinacantán Tzotzil, Malagasy, Dutch, Kanuri, White Hmong, Mapudungun, Hausa, Lower Sorbian
“High borrowers”
(25.9 – 45.6%)
Takia, Thai, Yaqui, Swahili, Vietnamese, Sakha, Archi, Imbabura Quechua, Kildin Saami, Bezhta, Indonesian, Japanese, Ceq Wong, Sarmaccan, English, Romanian, Gurindji
“Very high borrowers”
(51.7 – 62.7%)
Tarifyt Berber, Selice Romani

Another goal of the WOLD project was to learn more about borrowing in general. The research confirmed several generally accepted principles about borrowings:

  • Function words were borrowed less than content words (nouns, verbs, adjectives, and adverbs). Overall, 12% of function words were borrowed, compared to 25% of content words.
  • Nouns were more likely to be borrowed (31%) than other types of content words (14-15%).
  • Borrowing was most common for cultural vocabulary, such as religion, clothing, housing, law, social and political relations, agriculture, food, and warfare; and least common for personal vocabulary, such as sense perception, spatial relations, body parts, and kinship.

Motivation

My interest in the WOLD methodology dates from 2018, when I was starting to work on my second book, Bringing Linguistics into the Spanish Language Classroom. The book is organized around five themes, or “essential questions,” including “How is Spanish different from other languages?” and “How is Spanish similar to other languages?” I thought it would be interesting to compare Spanish to the WOLD languages so that I could say either “Spanish has borrowed more words than most other languages” or “Spanish has borrowed a typical amount of words.” (I was confident that Spanish would be a “low borrower.”)

I originally imagined that I could research this topic in a couple of weeks, but soon ran into methodological issues such as:

  • Should word pairs like hijo and hija (‘son/daughter’) be counted as two separate words, even though they are just masculine and feminine forms of the same word?
  • WOLD linguists could identify multiple words for a single meaning. How far should this be taken for Spanish? How does one draw the line between synonyms and dialectal variants?
  • When looking up word origins, the WOLD guidelines count a word as borrowed if it entered the language at any point in the language’s history. This would include, for instance, words borrowed into Classical or Vulgar Latin, such as gato ‘cat.’ (Vulgar Latin cattus is believed to be Afro-Asiatic in origin, and replaced the original Latin feles.) This guideline rubbed me the wrong way. Shouldn’t Spanish begin with Vulgar Latin?

After three months of a futile quick-and-dirty run at these issues, I decided to put the project on my back burner and eventually do a more thorough job that would hopefully yield publishable results. So…here we are.