
The WordWars Dataset
March 2020

Contact: 
Saif M. Mohammad (saif.mohammad@nrc-cnrc.gc.ca)

Project homepage: http://saifmohammad.com/WebPages/wordwars.html

This directory contains the WordWars data. Details of the dataset are in the 
paper below.

Papers

WordWars: A Dataset to Examine the Natural Selection of Words. Saif M. Mohammad.
In Proceedings of the 12th Language Resources and Evaluation Conference
(LREC-2020), Marseille, France.

There is a growing body of work on how word meaning changes over time: mutation.
In contrast, there is very little work on how different words compete to
represent the same meaning, and how the degree of success of words in that
competition changes over time: natural selection. We present a new dataset,
WordWars, with historical frequency data from the early 1800s to the early 2000s
for monosemous English words in over 5000 synsets. We explore three broad
questions with the dataset: (1) what is the degree to which predominant words in
these synsets have changed, (2) how do prominent word features such as
frequency, length, and concreteness impact natural selection, and (3) what are
the differences between the predominant words of the 2000s and the predominant
words of early 1800s. We show that close to one third of the synsets undergo a
change in the predominant word in this time period. Manual annotation of these
pairs shows that about 15% of these are orthographic variations, 25% involve
affix changes, and 60% have completely different roots. We find that frequency,
length, and concreteness all impact natural selection, albeit in different ways.

Keywords: Natural selection, lexical semantics, words, evolution, word length, frequency, concreteness


NOTES:

- The synsets are from WordNet 3.1. I have given my own unique integer ids for the synsets.
  These ids are not from WordNet 3.1.

- The term 'flip' in some of the files is short for 'change-of-winner-1809-2009'.

