Evidence before resemblance

Method & evidence

1 · The task: open-set ranking, not classification

StyleMatch treats authorship analysis as open-set profile ranking plus verification. A closed-set classifier must answer with one of its known authors; an open-set system must also be able to answer none of them. Every query is thresholded, and the system returns "no strong match" when the evidence is weak. Evaluation follows the same logic: complete sources — never random chunks — are held out, because chunks of one book on both sides of a split leak topic and edition artifacts and inflate accuracy (Sawatphol et al. 2024).

2 · Architecture: two physically separate channels

3 · Corpus discipline

Profiles are built exclusively from original-language primary texts. Translations, adaptations, subtitles, summaries, and generated imitations are excluded: translationese measurably shifts syntax, punctuation, register, and language-specific distinctions, so a translated profile partly measures the translator. Mirrors and editions of one work share an independent_source_id, so duplicated text can never count as independent evidence. Profiles without sufficient independent-source evidence remain retrievable but are excluded from headline evaluation. Rights-cleared in-copyright texts are indexed for matching but their passages are never displayed.

4 · How the deployed model was chosen

Six candidate views and a learned reranker were compared on identical source-heldout dev/test splits. This is a ranked-retrieval task, so MRR and Recall@3 are the primary measures; classification F1 would not measure whether the correct result rises through the shortlist.

Locked source-heldout MRR Point estimate · 95% bootstrap interval
Deployed Numerical leader
MethodDescription
Pretrained mStyleDistance A zero-shot test of how far a general multilingual style geometry transfers without project-specific authors or sources, separating native encoder strength from gains due to corpus adaptation.
Fine-tuned mStyleDistance The same architecture adapted with positive pairs from different independent works, language-aware batches, and matched hard negatives that discourage language, topic, register, and source shortcuts.
mStyleDistance with source prototypes A work-level representation that retains a prototype for each independent source, testing whether variation across books, speeches, and periods improves retrieval of genuinely unseen material.
Pretrained multilingual authorship representation An authorship-oriented multilingual encoder evaluated without StyleMatch adaptation to measure direct transfer into the project’s literary and rhetorical corpus.
Fine-tuned multilingual authorship representation The authorship encoder trained under the source-separated protocol, requiring same-author evidence to cross independent works; this representation supplies the deployed author-profile ranking.
Classical style-feature fusion An interpretable non-topic baseline combining delexicalized character patterns, function words, punctuation, sentence rhythm, discourse markers, and compression distance.
Learned multi-view reranker A candidate reranker trained on within-language normalized neural and classical scores, testing whether disagreement among views adds stable evidence under bootstrap and subgroup non-degradation gates.

Pre-declared selection rule. The reranker led by .003 MRR, but its paired-bootstrap interval for the difference crossed zero [−.052, .060] and the subgroup non-degradation gate was not met. The best dev-selected single model therefore remains deployed.

Expanded-candidate stress test. A source-grouped, cross-fitted density penalty reduced false top-three concentration (HHI .00546 → .00522; Gini .240 → .204), while its paired MRR and Recall@3 intervals included no change. It therefore remains an exploratory audit, not a production ranking term (Radovanović et al. 2010; Schnitzer et al. 2012).

5 · Open-set calibration

For each language, held-out authors are split into "known" and "unknown"; a logistic calibrator is fitted on the maximum profile similarity, and the equal-error-rate threshold becomes the rejection bar. A calibrator is accepted only when it was fitted with the exact encoder revision used by the index. Confidence labels therefore retain a direct, language-specific link to the model and evidence that produced them.

6 · Performance record

Current system facts and warm-query reference timings are listed below; each measurement keeps its evaluation scope visible.

System propertyCurrent recordWhy it matters
Multilingual coverage 9 languages One shared retrieval system with language-aware result grouping
Evidence memory 2,577 source prototypes · 63,576 chunks Independent works preserve variation beyond one averaged voice
Selected style model Fine-tuned multilingual authorship representation The evaluated winner and the served encoder are the same method
Retrieval engine Exact normalized inner product Stable deterministic ranking across repeated queries
Warm query latency — CPU p50 119.2 ms · p95 124.0 ms 30-run within-language reference benchmark
Warm query latency — GPU p50 25.7 ms · p95 26.9 ms 30-run within-language reference benchmark
Score contract Style · Topic/Tone · Affinity shown separately A user can see which signal produced each part of the result
Text policy Direct original-language comparison Translation does not overwrite authorial syntax or register

7 · Known limitations

8 · References

  1. Kittler, Hatef, Duin & Matas (1998). On Combining Classifiers. IEEE TPAMI.
  2. Montague & Aslam (2001). Relevance Score Normalization for Metasearch. CIKM.
  3. Cameron, Gelbach & Miller (2008). Bootstrap-Based Improvements for Inference with Clustered Errors. Review of Economics and Statistics.
  4. Cawley & Talbot (2010). On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation. JMLR.
  5. Radovanović, Nanopoulos & Ivanović (2010). Hubs in Space: Popular Nearest Neighbors in High-Dimensional Data. JMLR.
  6. Schnitzer et al. (2012). Local and Global Scaling Reduce Hubs in Space. JMLR.
  7. Guo, Pleiss, Sun & Weinberger (2017). On Calibration of Modern Neural Networks. ICML.
  8. Geifman & El-Yaniv (2017). Selective Classification for Deep Neural Networks. NeurIPS.
  9. Dror et al. (2018). The Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing. ACL.
  10. Sagawa, Koh, Hashimoto & Liang (2020). Distributionally Robust Neural Networks for Group Shifts. ICLR.
  11. Sawatphol, Udomcharoenchaikit & Nutanong (2024). Addressing Topic Leakage in Cross-Topic Evaluation for Authorship Verification. TACL.
  12. Huertas-Tato et al. (2024). Isolating authorship from content with semantic embeddings and contrastive learning. arXiv:2411.18472.
  13. Terreau, Gourru & Velcin (2024). Capturing Style in Author and Document Representation. arXiv:2407.13358.
  14. Wang et al. (2024). Multilingual E5 Text Embeddings: A Technical Report. arXiv:2402.05672.
  15. Qiu et al. (2025). mStyleDistance: Multilingual Style Embeddings and their Evaluation. Findings of ACL.
  16. Kim, Zhang & Jurgens (2025). Leveraging Multilingual Training for Authorship Representation. EMNLP.
  17. Alshomary et al. (2025). Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer Layers. EMNLP.
  18. Alshomary et al. (2025). Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution. COLING.
  19. Alipoormolabashi, Patel & Balasubramanian (2025). Quantifying Misattribution Unfairness in Authorship Attribution. ACL.
  20. Icard et al. (2025). Embedding Style Beyond Topics: Analyzing Dispersion Effects Across Different Language Models. COLING.
  21. Man et al. (2026). Explainable Disentangled Representation Learning for Generalizable Authorship Attribution. ACL.
  22. Anand, Alshomary & McKeown (2026). iBERT: Interpretable Embeddings via Sense Decomposition. EACL.