Editors’ Choice: Treating Texts as Individuals vs. Lumping Them Together
Ted Underwood has been talking up the advantages of the Mann-Whitney test over Dunning’s Log-likelihood which is currently more widely used. I’m having trouble getting M-W running on large numbers of texts as quickly as I’d like, but I’d say that his basic contention–that Dunning log-likelihood is frequently not the best method–is definitely true, and there’s a lot to like about rank-ordering tests.
Before I say anything about the specifics, though, I want to make a more general point first, about how we think about comparing groups of texts.The most important difference between these two tests rests on a much bigger question about how to treat the two corpuses we want to compare.