WordNet Domains (XWND)

This page documents how subject-domain labels from eXtended WordNet Domains (XWND) can be used with the Open Multilingual Wordnet. Domain labels are mapped to the Interlingual Index (ILI) so that they apply across all wordnets that share ILI entries.

Background

WordNet Domains (WND) (Magnini & Cavaglià, 2000) is a manually constructed resource that assigns each Princeton WordNet synset one or more labels from a hierarchy of 170 subject-domain labels derived from the Dewey Decimal Classification (e.g., medicine, sport, gastronomy). The 170 domains are organized into a hierarchy: sport is a subdomain of factotum, gastronomy of food, and so on.

eXtended WordNet Domains (XWND) (González, Rigau & Castillo, 2012) propagates these labels to all WordNet 3.0 synsets using the UKB graph-based algorithm over the WordNet relation graph enriched with gloss relations, and provides a confidence score for each domain assignment. This gives 82,115 noun synsets a primary domain label, compared to partial coverage in the original WND.

We assign each synset its highest-scoring domain as its primary label. The XWND data and this mapping were used in the analysis of metaphor and metonymy reported in the JPC1 project.

Why domains differ from topics

WordNet topics (lexicographer files / supersenses) are editorial labels assigned by the WordNet team to organise their work — they group polysemous senses that share a broad field. XWND domains, by contrast, are derived from the Dewey Decimal Classification and capture subject matter rather than lexicographic structure. The two systems partition the lexicon very differently. For example, the topic noun.cognition covers all mental-state nouns regardless of their subject domain, while the XWND domain psychology groups nouns by their academic field.

An interactive treemap of all 170 domains — sized by synset count and coloured by top-level branch — is available on a separate page. Click any rectangle to zoom in.

Download

A precomputed tab-separated file mapping ILI to highest-weight XWND domain is available here:

ili_domains.tsv (ILI → domain, 82,115 noun synsets, tab-separated with header)

Format:

ILI	domain
i100000	plants
i100001	biology
i100002	plants
…

Generating the mapping yourself

The script below downloads the XWND data from the JPC1 repository, loads omw-en:1.4 via the wn Python library, selects the highest-scoring domain per synset, and writes the result as a tab-separated file. Run it with uv run get_domains.py (no separate install step needed).

#!/usr/bin/env python3
# /// script
# requires-python = ">=3.10"
# dependencies = [
#   "wn>=1.1",
# ]
# ///
"""Map XWND domains to ILI and write a tab-separated file."""

import tarfile
import urllib.request
from pathlib import Path
import wn

XWND_URL = "https://github.com/bond-lab/JPC1/raw/main/tasks/wordnet/xwnd-30g.tgz"
XWND_DIR = Path("xwnd-30g")
OUT_TSV  = Path("ili_domains.tsv")

# --- download and extract XWND if needed ---
if not XWND_DIR.exists():
    print("Downloading XWND data...")
    urllib.request.urlretrieve(XWND_URL, "xwnd-30g.tgz")
    with tarfile.open("xwnd-30g.tgz") as tar:
        tar.extractall(".")
    print("Extracted.")

# --- load highest-weight domain per synset from .ppv files ---
best: dict[str, tuple[str, float]] = {}   # offset-pos -> (domain, weight)
for ppv_path in sorted(XWND_DIR.glob("*.ppv")):
    domain = ppv_path.stem
    with open(ppv_path) as f:
        for line in f:
            parts = line.split()
            if len(parts) != 2:
                continue
            key, weight_str = parts
            if not key.endswith("-n"):
                continue
            weight = float(weight_str)
            prev = best.get(key)
            if prev is None or weight > prev[1]:
                best[key] = (domain, weight)

# offset-pos key (e.g. "00001740-n") -> omw-en synset ID
domains = {f"omw-en-{key}": dom for key, (dom, _) in best.items()}

# --- map to ILI via the wn library ---
wn.download("omw-en:1.4")           # skipped if already downloaded
my_wn = wn.Wordnet(lexicon="omw-en:1.4")

rows = []
for ss in my_wn.synsets(pos="n"):
    domain = domains.get(ss.id)
    if domain and ss.ili:
        rows.append((ss.ili, domain))
rows.sort()

with open(OUT_TSV, "w") as f:
    f.write("ILI\tdomain\n")
    for ili, domain in rows:
        f.write(f"{ili}\t{domain}\n")

print(f"Wrote {len(rows):,} rows to {OUT_TSV}")

Using the mapping in Python

Once you have ili_domains.tsv (or generated it with the script above), you can use it with any wordnet that provides ILI entries:

import wn

# Load the mapping
domains: dict[str, str] = {}
with open("ili_domains.tsv") as f:
    next(f)  # skip header
    for line in f:
        ili, domain = line.rstrip("\n").split("\t")
        domains[ili] = domain

# Look up the domain for a word
my_wn = wn.Wordnet(lexicon="omw-en:1.4")
for ss in my_wn.words("dog", pos="n")[0].synsets():
    print(ss.ili, domains.get(ss.ili, "—"), ss.definitions()[0])

References


Maintainer: Francis Bond

Source code hosted at https://github.com/omwn/omwn.github.io.