Puzzle Book Smith

Data sources

Last revised: 4 October 2026

About this page

Puzzle Book Smith builds its word lists and checks from open data. This page names each source, says what we use it for, and gives its licence. Thank you to everyone who shares this work.

wordfreq (word frequency)

Word frequency data: wordfreq by Robyn Speer, licensed under CC BY-SA 4.0. It includes SUBTLEX data by Marc Brysbaert and colleagues, and Google Books Ngram data (books.google.com/ngrams).

We use it to tell everyday words from rare ones. That keeps puzzle words familiar. Our everyday word list is derived from wordfreq and is shared under CC BY-SA 4.0.

What we changed: we kept words that are common in English, limited them to 3 to 12 letters, and removed names, unpleasant words and obscure words using our own lists.

SCOWL (word lists)

Word lists: SCOWL, copyright 2000-2026 Kevin Atkinson. It is free to use, copy and modify when the copyright notice is kept.

We use it as a base spelling list and a size band for each word. Words it flags as offensive are left out of our family safe word bank.

Open English WordNet (word relations)

Compound word checks use Open English WordNet 2025 by John P. McCrae and the Open English WordNet contributors, licensed under CC BY 4.0. It is based on Princeton WordNet 3.1 (WordNet License).

We use it to confirm that a compound such as "sun" and "flower" is a real word or phrase before it appears in a Word Match puzzle. We filtered the list and reviewed it by hand.

cuss (offensive word list)

cuss, copyright 2016 Titus Wormer, MIT licence.

We use it to screen offensive words out of our family safe word bank.