Learn how outlier detection depends on the text representation and why length belongs beside distance.
Owen has one shelf for speeches that deserve closer reading. He asks for the addresses most unlike the rest, expecting a neutral ranking.
The ranking arrives with a catch. An outlier is a case that looks far from other cases under a chosen measurement. For text, changing the features can change what looks far away, and very short documents can look strange because they have little vocabulary to average.
TipWhat you will learn
This lesson shows how to:
build speech-level document-feature matrices;
compute cosine distance from the average speech;
report word length next to distance;
test how much distance tracks length; and
compare rank agreement with a shuffled-feature null.
Build the representations
A representation is the version of the text that a method sees. Here one representation keeps content words after removing stopwords, another gives those content words tf-idf weights, and a third keeps all words as raw counts. Cosine distance compares the direction of two word-count vectors. A value closer to 1 means farther from the average vector.
Most and least distant speeches using content-word raw counts
Group
Speech
Year
President
Total words
Content words
Shortest in corpus
Cosine distance
Most distant
1793-Washington
1793
George Washington
135
62
yes
0.677
Most distant
2021-Biden
2021
Joseph R. Biden
407
188
no
0.639
Most distant
1865-Lincoln
1865
Abraham Lincoln
698
338
no
0.605
Most distant
2017-Trump
2017
Donald J. Trump
713
346
no
0.601
Most distant
1945-Roosevelt
1945
Franklin D. Roosevelt
511
238
no
0.600
Least distant
1897-McKinley
1897
William McKinley
3960
1931
no
0.235
Least distant
1841-Harrison
1841
William Henry Harrison
8446
3796
no
0.247
Least distant
1925-Coolidge
1925
Calvin Coolidge
4053
1879
no
0.255
Least distant
1881-Garfield
1881
James A. Garfield
2951
1430
no
0.263
Least distant
1845-Polk
1845
James Knox Polk
4802
2263
no
0.276
The 1793-Washington speech is the most distant under this representation, and it is the shortest speech at 135 total words. A 135-word document has little vocabulary to average, so a handful of words can drive its position.
Check the length confound
A confound is a second factor that can explain a result the analyst might otherwise attribute to the method’s target. Here the candidate confound is document length.
tfidf_rank <- lengths |>mutate(distance =cosine_from_average(content_tfidf_dfm),representation ="content words, tf-idf" ) |>arrange(desc(distance))all_words_rank <- lengths |>mutate(distance =cosine_from_average(all_words_dfm),representation ="all words, raw counts" ) |>arrange(desc(distance))all_rankings <-bind_rows(content_rank, tfidf_rank, all_words_rank)length_correlations <- all_rankings |>summarise(correlation =round(cor(total_words, distance), 3),.by = representation )length_threshold <-1000Lrestricted_top <- all_rankings |>filter(total_words >= length_threshold) |>arrange(desc(distance), .by_group =TRUE) |>slice_head(n =5, by = representation) |>transmute( representation, speech_id, year, total_words,distance =round(distance, 3) )kable( length_correlations,col.names =c("Representation", "Correlation between total words and distance"),caption ="Distance is strongly related to speech length",row.names =FALSE)
Distance is strongly related to speech length
Representation
Correlation between total words and distance
content words, raw counts
-0.738
content words, tf-idf
-0.908
all words, raw counts
-0.534
kable( restricted_top,col.names =c("Representation", "Speech", "Year", "Total words", "Cosine distance"),caption ="Top outliers after restricting to speeches with at least 1,000 words",row.names =FALSE)
Top outliers after restricting to speeches with at least 1,000 words
Representation
Speech
Year
Total words
Cosine distance
content words, tf-idf
1813-Madison
1813
1152
0.805
content words, tf-idf
1965-Johnson
1965
1362
0.782
content words, tf-idf
1941-Roosevelt
1941
1118
0.781
content words, tf-idf
1977-Carter
1977
1185
0.778
content words, tf-idf
1789-Washington
1789
1420
0.769
content words, raw counts
1813-Madison
1813
1152
0.554
content words, raw counts
1809-Madison
1809
1175
0.503
content words, raw counts
1941-Roosevelt
1941
1118
0.494
content words, raw counts
1961-Kennedy
1961
1320
0.471
content words, raw counts
1849-Taylor
1849
1088
0.451
all words, raw counts
2001-Bush
2001
1289
0.146
all words, raw counts
1977-Carter
1977
1185
0.128
all words, raw counts
2025-Trump
2025
2555
0.124
all words, raw counts
1993-Clinton
1993
1568
0.115
all words, raw counts
1973-Nixon
1973
1491
0.112
The correlations are negative under all three representations: -0.738 for content-word counts, -0.908 for content-word tf-idf, and -0.534 for all-word counts. Shorter speeches tend to sit farther from the average. Once speeches under 1,000 words are removed, 1793-Washington and 2021-Biden leave the top lists. The outlier ranking is substantially a length ranking.
Compare the feature choices with a null
Now add tf-idf weights and then keep stopwords in a separate raw-count representation. This changes both the weighting rule and the feature set. A rank agreement statistic checks whether the three distance rankings put speeches in similar order.
A permutation null reruns the rank-agreement calculation after feature counts are randomized. The p-value is the fraction of shuffled agreements that are at least as large as the observed agreement.
kable( rank_agreement_display,col.names =c("Observed", "Null mean", "Null low", "Null high", "p-value", "Replicates", "Alternative"),caption ="Permutation null for rank agreement across representations",row.names =FALSE)
Permutation null for rank agreement across representations
Observed
Null mean
Null low
Null high
p-value
Replicates
Alternative
0.77
0.953
0.939
0.966
1
500
greater
Content-word raw counts and content-word tf-idf share all five top speeches. Keeping stopwords shares four of five with content-word raw counts and changes the third slot from 1865-Lincoln to 2001-Bush. The average Spearman rank agreement across the three full rankings is 0.770. When document lengths are held fixed but features are randomly reassigned, the null mean is 0.953 with a 90% interval from 0.939 to 0.966. The observed agreement does not exceed that shuffled-feature null (p = 1). Cross-representation agreement is therefore not reassuring by itself.
Outlier methods always rank something first. The shuffle removes document-specific word patterns while preserving lengths, so it asks whether the three representations agree more than feature noise would. They do not. The shortest-speech result remains a length-confounded reading lead.
What to remember
Cosine distance ranks each speech against the average vector.
The content-word representations have 9,193 features; the all-word representation has 9,327.
Under content-word raw counts, 1793-Washington is farthest and 1897-McKinley is closest to the average.
Distance is strongly tied to length here; the shortest speeches drive the top of the list.
The three representation rankings do not agree more than shuffled features predict after lengths are fixed.
Use outlier scores to choose what to read next. Do not treat the ranking as a property of the speech apart from the features you chose.