Learn how speech-level document similarity changes under raw counts, tf-idf, and stopword removal.
A reading group wants one companion speech for a long weekend assignment. They ask for the nearest neighbour to a selected address.
The request sounds simple until the group asks what “nearest” means. A method can compare word counts, rare-word weights, or a version with common words removed.
Document similarity is a score that compares two documents after they have been turned into features. The score is only as meaningful as that representation for the question at hand.
TipWhat you will learn
This lesson shows how to:
build speech-level document-feature matrices;
compute cosine similarity across 60 speeches;
list the most and least similar pairs under tf-idf;
show that nearest neighbours change when features change;
test an authorship pattern against shuffled labels; and
notice when document length is part of the signal.
Build speech vectors
Tf-idf stands for term frequency times inverse document frequency. It keeps words from each speech, then gives more weight to words that are less common across the full set of speeches.
Three speech representations used for similarity comparisons
Representation
Documents
Features
raw counts
60
9327
tf-idf
60
9327
tf-idf with stopwords removed
60
9193
kable( retention_by_era |>mutate(share_retained =round(share_retained, 3)),col.names =c("Era", "Speeches", "Original words", "Retained words", "Share retained", "Median kept paragraph words"),caption ="Words retained after rebuilding speeches from paragraphs of at least 25 words",row.names =FALSE)
Words retained after rebuilding speeches from paragraphs of at least 25 words
Era
Speeches
Original words
Retained words
Share retained
Median kept paragraph words
before 1900
28
71961
71800
0.998
112
1900 or later
32
69032
63106
0.914
54
kable( identity_table,col.names =c("Item", "Count"),caption ="The corpus has 60 speeches by 40 people but only 36 surnames",row.names =FALSE)
The corpus has 60 speeches by 40 people but only 36 surnames
Item
Count
speeches
60
people
40
surnames
36
surnames used by two people
4
The helper rebuilds each speech by joining only paragraphs with at least 25 words, so the “speeches” below are retained-paragraph versions of the addresses. The filter keeps nearly all words before 1900 and a smaller share from 1900 or later. The president field is the full name: 60 speeches come from 40 people, while four surnames each refer to two people.
Comparing every speech with every other speech gives 1,770 pairs, because 60 * 59 / 2 = 1,770.
Rank all speech pairs
Cosine similarity ranges from 0 to 1 for these nonnegative vectors. Larger values mean the two speeches point in more similar vocabulary directions.
Most and least similar speech pairs under speech-level tf-idf
Group
First speech
Second speech
First president
Second president
First tokens
Second tokens
Cosine similarity
Most similar
1817-Monroe
1821-Monroe
James Monroe
James Monroe
3373
4470
0.2903
Most similar
1837-VanBuren
1841-Harrison
Martin Van Buren
William Henry Harrison
3846
8465
0.2471
Most similar
1897-McKinley
1909-Taft
William McKinley
William Howard Taft
3973
5430
0.2454
Most similar
1841-Harrison
1845-Polk
William Henry Harrison
James Knox Polk
8465
4810
0.2438
Most similar
1825-Adams
1845-Polk
John Quincy Adams
James Knox Polk
2920
4810
0.2365
Least similar
1793-Washington
1905-Roosevelt
George Washington
Theodore Roosevelt
135
984
0.0066
Least similar
1793-Washington
2021-Biden
George Washington
Joseph R. Biden
135
412
0.0073
Least similar
1793-Washington
1941-Roosevelt
George Washington
Franklin D. Roosevelt
135
1143
0.0074
Least similar
1829-Jackson
2021-Biden
Andrew Jackson
Joseph R. Biden
1130
412
0.0075
Least similar
1793-Washington
1913-Wilson
George Washington
Woodrow Wilson
135
1699
0.0095
The strongest pair under this representation is the two Monroe speeches, with cosine similarity about 0.2903. Four of the five weakest rows include the short 1793 Washington speech, which has 135 retained tokens. Shortness, not a reading of subject matter, puts that speech at the bottom of the ranking.
Mean tf-idf similarity compared with speech token count
Measure
Correlation
Mean tf-idf similarity and speech token count
0.861
Mean tf-idf similarity still correlates strongly with speech length. Tf-idf changes the representation, but it does not remove the length gradient from this corpus.
Change the representation
A nearest neighbour is the highest-scoring other document. The code below asks for the nearest neighbour of the 1793 Washington speech three ways: raw counts, tf-idf, and tf-idf after removing stopwords.
One target speech gets three nearest neighbours under three representations
Target
Representation
Nearest neighbour
Year
President
Neighbour tokens
Cosine similarity
1793-Washington
raw counts
1841-Harrison
1841
William Henry Harrison
8465
0.8474
1793-Washington
tf-idf
1861-Lincoln
1861
Abraham Lincoln
3617
0.0552
1793-Washington
tf-idf, stopwords removed
1885-Cleveland
1885
Grover Cleveland
1687
0.0512
The three answers differ. The raw-count neighbour for the shortest speech is the longest retained speech, and the score is about 0.8474. That column is dominated by shared high-frequency words and length, so it should not be read as a topical match.
How often each speech is the nearest neighbour under raw counts
Raw-count nearest neighbour
Number of target speeches
1985-Reagan
7
1925-Coolidge
6
1897-McKinley
4
1953-Eisenhower
4
1845-Polk
3
1853-Pierce
3
1877-Hayes
3
1889-Harrison
3
1909-Taft
3
1817-Monroe
2
The raw-count hub check does not make 1841-Harrison a corpus-wide hub; it is the nearest neighbour for two speeches. The broader warning is still clear. Several speeches become nearest neighbours for many targets under raw counts, which is another sign that the column measures common-word mass more than a careful vocabulary match.
Separate vocabulary signal from authorship signal
When one person has more than one address in the corpus, repeated phrasing can make those speeches neighbours. That can be useful for authorship or style questions and distracting for topic questions.
Same-president tf-idf similarity compared with shuffled author labels
Observed
Null mean
Null 5th pct
Null 95th pct
p-value
Replicates
Alternative
0.128
0.081
0.068
0.095
0.001
1000
greater
Under tf-idf, 19 of the 60 speeches have a nearest neighbour by the same person. The label-shuffle null says same-person pairs are more similar than chance in this corpus. That is evidence of repeated wording or style, not proof that the speeches discuss the same subjects.
What to remember
Document similarity compares representations, not documents in the abstract.
Speech-level tf-idf gives 1,770 pairwise comparisons for 60 speeches.
The nearest neighbour of one speech changed across raw counts, tf-idf, and stopword removal.
Same-person neighbours can show authorship or style rather than subject matter.
Similarity functions return a ranking even when length or labels explain it.
The reading group chooses a companion speech only after naming what kind of similarity it wants.