Concept Inclusiveness

Inclusiveness measures how broadly a concept generalises within the WordNet noun hierarchy. It was introduced by Iliev & Axelrod (2017) to study the precision versus concreteness trade-off in language, and provides a finer-grained view of generality than raw depth in the hierarchy.

Definition

For a noun synset c that has n unique descendants (both direct and indirect hyponyms) among N total noun synsets in the wordnet:

inclusiveness(c) = −log ((n + 1) / N)

The “+1” avoids a zero in the numerator for leaf nodes (concepts with no hyponyms). The negation of the logarithm means:

Inclusiveness differs from hierarchy depth because depth measures how far down a concept sits, while inclusiveness measures how much of the hierarchy it covers. A concept at depth 5 may still be highly inclusive if it has thousands of descendants.

Computing inclusiveness

The script below computes inclusiveness for all noun synsets in any wordnet loaded via the wn Python library. Run it with uv run inclusiveness.py.

#!/usr/bin/env python3
# /// script
# requires-python = ">=3.10"
# dependencies = [
#   "wn>=1.1",
# ]
# ///
"""Compute concept inclusiveness for all noun synsets in a wordnet."""

import math
import wn

wn.download("omw-en:1.4")   # skipped if already present
my_wn = wn.Wordnet(lexicon="omw-en:1.4")

noun_synsets = my_wn.synsets(pos="n")

# Build hyponym adjacency list
str_hyponyms: dict[str, list[str]] = {
    ss.id: [h.id for h in ss.hyponyms()] for ss in noun_synsets
}
ids = list(str_hyponyms.keys())
idx = {ss_id: i for i, ss_id in enumerate(ids)}
children = [
    [idx[h] for h in str_hyponyms[ss_id] if h in idx] for ss_id in ids
]
N = len(ids)

# Count descendants via BFS/DFS for each synset
def count_descendants(i: int) -> int:
    visited: set[int] = set()
    stack = children[i][:]
    while stack:
        j = stack.pop()
        if j not in visited:
            visited.add(j)
            stack.extend(children[j])
    return len(visited)

inclusiveness = {
    ids[i]: -math.log((count_descendants(i) + 1) / N)
    for i in range(N)
}

# Example output: look up a word
for ss in my_wn.words("dog", pos="n")[0].synsets():
    score = inclusiveness.get(ss.id)
    if score is not None:
        print(f"{ss.id}  inclusiveness={score:.3f}  {ss.definitions()[0]}")

The computation visits each synset's descendant subgraph, so it takes a few minutes for large wordnets (~80,000 noun synsets). Cache the result as JSON if you need it repeatedly:

import json, math
from pathlib import Path

cache = Path("inclusiveness.json")
if cache.exists():
    inclusiveness = json.loads(cache.read_text())
else:
    # ... run the computation above ...
    cache.write_text(json.dumps(inclusiveness))

Relationship to Information Content

Inclusiveness is related to, but distinct from, Information Content (IC). IC is corpus-based: it measures how rarely a synset (and its hyponyms) appears in a large corpus, using the formula −log(P(c)). Inclusiveness is purely structural: it counts hyponym coverage in the ontology, without any corpus. The two measures are correlated — rare, specific concepts tend to appear less in corpora — but inclusiveness has the advantage of requiring no corpus and being computable for any wordnet regardless of language.

References


Maintainer: Francis Bond

Source code hosted at https://github.com/omwn/omwn.github.io.