This page documents how subject-domain labels from eXtended WordNet Domains (XWND) can be used with the Open Multilingual Wordnet. Domain labels are mapped to the Interlingual Index (ILI) so that they apply across all wordnets that share ILI entries.
WordNet Domains (WND) (Magnini & Cavaglià, 2000) is a manually constructed resource that assigns each Princeton WordNet synset one or more labels from a hierarchy of 170 subject-domain labels derived from the Dewey Decimal Classification (e.g., medicine, sport, gastronomy). The 170 domains are organized into a hierarchy: sport is a subdomain of factotum, gastronomy of food, and so on.
eXtended WordNet Domains (XWND) (González, Rigau & Castillo, 2012) propagates these labels to all WordNet 3.0 synsets using the UKB graph-based algorithm over the WordNet relation graph enriched with gloss relations, and provides a confidence score for each domain assignment. This gives 82,115 noun synsets a primary domain label, compared to partial coverage in the original WND.
We assign each synset its highest-scoring domain as its primary label. The XWND data and this mapping were used in the analysis of metaphor and metonymy reported in the JPC1 project.
WordNet topics (lexicographer files / supersenses) are editorial labels assigned by the WordNet team to organise their work — they group polysemous senses that share a broad field. XWND domains, by contrast, are derived from the Dewey Decimal Classification and capture subject matter rather than lexicographic structure. The two systems partition the lexicon very differently. For example, the topic noun.cognition covers all mental-state nouns regardless of their subject domain, while the XWND domain psychology groups nouns by their academic field.
An interactive treemap of all 170 domains — sized by synset count and coloured by top-level branch — is available on a separate page. Click any rectangle to zoom in.
A precomputed tab-separated file mapping ILI to highest-weight XWND domain is available here:
ili_domains.tsv (ILI → domain, 82,115 noun synsets, tab-separated with header)
Format:
ILI domain i100000 plants i100001 biology i100002 plants …
The script below downloads the XWND data from the JPC1 repository, loads omw-en:1.4 via the wn Python library, selects the highest-scoring domain per synset, and writes the result as a tab-separated file. Run it with uv run get_domains.py (no separate install step needed).
#!/usr/bin/env python3
# /// script
# requires-python = ">=3.10"
# dependencies = [
# "wn>=1.1",
# ]
# ///
"""Map XWND domains to ILI and write a tab-separated file."""
import tarfile
import urllib.request
from pathlib import Path
import wn
XWND_URL = "https://github.com/bond-lab/JPC1/raw/main/tasks/wordnet/xwnd-30g.tgz"
XWND_DIR = Path("xwnd-30g")
OUT_TSV = Path("ili_domains.tsv")
# --- download and extract XWND if needed ---
if not XWND_DIR.exists():
print("Downloading XWND data...")
urllib.request.urlretrieve(XWND_URL, "xwnd-30g.tgz")
with tarfile.open("xwnd-30g.tgz") as tar:
tar.extractall(".")
print("Extracted.")
# --- load highest-weight domain per synset from .ppv files ---
best: dict[str, tuple[str, float]] = {} # offset-pos -> (domain, weight)
for ppv_path in sorted(XWND_DIR.glob("*.ppv")):
domain = ppv_path.stem
with open(ppv_path) as f:
for line in f:
parts = line.split()
if len(parts) != 2:
continue
key, weight_str = parts
if not key.endswith("-n"):
continue
weight = float(weight_str)
prev = best.get(key)
if prev is None or weight > prev[1]:
best[key] = (domain, weight)
# offset-pos key (e.g. "00001740-n") -> omw-en synset ID
domains = {f"omw-en-{key}": dom for key, (dom, _) in best.items()}
# --- map to ILI via the wn library ---
wn.download("omw-en:1.4") # skipped if already downloaded
my_wn = wn.Wordnet(lexicon="omw-en:1.4")
rows = []
for ss in my_wn.synsets(pos="n"):
domain = domains.get(ss.id)
if domain and ss.ili:
rows.append((ss.ili, domain))
rows.sort()
with open(OUT_TSV, "w") as f:
f.write("ILI\tdomain\n")
for ili, domain in rows:
f.write(f"{ili}\t{domain}\n")
print(f"Wrote {len(rows):,} rows to {OUT_TSV}")
Once you have ili_domains.tsv (or generated it with the script above), you can use it with any wordnet that provides ILI entries:
import wn
# Load the mapping
domains: dict[str, str] = {}
with open("ili_domains.tsv") as f:
next(f) # skip header
for line in f:
ili, domain = line.rstrip("\n").split("\t")
domains[ili] = domain
# Look up the domain for a word
my_wn = wn.Wordnet(lexicon="omw-en:1.4")
for ss in my_wn.words("dog", pos="n")[0].synsets():
print(ss.ili, domains.get(ss.ili, "—"), ss.definitions()[0])
Source code hosted at https://github.com/omwn/omwn.github.io.