Method & evidence
1 · The task: open-set ranking, not classification
StyleMatch treats authorship analysis as open-set profile ranking plus verification. A closed-set classifier must answer with one of its known authors; an open-set system must also be able to answer none of them. Every query is thresholded, and the system returns "no strong match" when the evidence is weak. Evaluation follows the same logic: complete sources — never random chunks — are held out, because chunks of one book on both sides of a split leak topic and edition artifacts and inflate accuracy (Sawatphol et al. 2024).
2 · Architecture: two physically separate channels
- Style channel. A multilingual authorship-representation encoder (Kim, Zhang & Jurgens 2025), contrastively fine-tuned on same-author pairs drawn from different works — so the objective cannot be solved by recognizing the book — with matched hard negatives and language-aware batches.
- Topic channel. A separate multilingual E5 encoder (Wang et al. 2024) used only for the topic score. Style and topic never share an embedding, an index, or a score.
- Profiles. Each author-language profile is a normalized centroid over 75–150-word original-language chunks, plus per-work source prototypes that preserve within-author variation. Retrieval is exact inner product over cached embeddings — no approximate search, no model calls per candidate.
- Composite. The displayed Affinity is a declared product rule (0.7·Style + 0.3·Topic within language; 0.5/0.5 across languages), not a learned or calibrated quantity. Sub-scores are always shown separately.
- Evidence. Matches carry interpretable stylometric observations (function words, punctuation, sentence rhythm, lexical diversity) computed against the candidate cohort — a complement to the encoder, never a replacement (Terreau et al. 2024).
3 · Corpus discipline
Profiles are built exclusively from original-language primary texts.
Translations, adaptations, subtitles, summaries, and generated imitations are excluded:
translationese measurably shifts syntax, punctuation, register, and language-specific
distinctions, so a translated profile partly measures the translator. Mirrors and
editions of one work share an independent_source_id, so duplicated text can
never count as independent evidence. Profiles without sufficient independent-source
evidence remain retrievable but are excluded from headline evaluation.
Rights-cleared in-copyright texts are indexed for matching but their passages are never
displayed.
4 · How the deployed model was chosen
Six candidate views and a learned reranker were compared on identical source-heldout dev/test splits. This is a ranked-retrieval task, so MRR and Recall@3 are the primary measures; classification F1 would not measure whether the correct result rises through the shortlist.
| Method | Description |
|---|---|
| Pretrained mStyleDistance | A zero-shot test of how far a general multilingual style geometry transfers without project-specific authors or sources, separating native encoder strength from gains due to corpus adaptation. |
| Fine-tuned mStyleDistance | The same architecture adapted with positive pairs from different independent works, language-aware batches, and matched hard negatives that discourage language, topic, register, and source shortcuts. |
| mStyleDistance with source prototypes | A work-level representation that retains a prototype for each independent source, testing whether variation across books, speeches, and periods improves retrieval of genuinely unseen material. |
| Pretrained multilingual authorship representation | An authorship-oriented multilingual encoder evaluated without StyleMatch adaptation to measure direct transfer into the project’s literary and rhetorical corpus. |
| Fine-tuned multilingual authorship representation | The authorship encoder trained under the source-separated protocol, requiring same-author evidence to cross independent works; this representation supplies the deployed author-profile ranking. |
| Classical style-feature fusion | An interpretable non-topic baseline combining delexicalized character patterns, function words, punctuation, sentence rhythm, discourse markers, and compression distance. |
| Learned multi-view reranker | A candidate reranker trained on within-language normalized neural and classical scores, testing whether disagreement among views adds stable evidence under bootstrap and subgroup non-degradation gates. |
Pre-declared selection rule. The reranker led by .003 MRR, but its paired-bootstrap interval for the difference crossed zero [−.052, .060] and the subgroup non-degradation gate was not met. The best dev-selected single model therefore remains deployed.
Expanded-candidate stress test. A source-grouped, cross-fitted density penalty reduced false top-three concentration (HHI .00546 → .00522; Gini .240 → .204), while its paired MRR and Recall@3 intervals included no change. It therefore remains an exploratory audit, not a production ranking term (Radovanović et al. 2010; Schnitzer et al. 2012).
5 · Open-set calibration
For each language, held-out authors are split into "known" and "unknown"; a logistic calibrator is fitted on the maximum profile similarity, and the equal-error-rate threshold becomes the rejection bar. A calibrator is accepted only when it was fitted with the exact encoder revision used by the index. Confidence labels therefore retain a direct, language-specific link to the model and evidence that produced them.
6 · Performance record
Current system facts and warm-query reference timings are listed below; each measurement keeps its evaluation scope visible.
| System property | Current record | Why it matters |
|---|---|---|
| Multilingual coverage | 9 languages | One shared retrieval system with language-aware result grouping |
| Evidence memory | 2,577 source prototypes · 63,576 chunks | Independent works preserve variation beyond one averaged voice |
| Selected style model | Fine-tuned multilingual authorship representation | The evaluated winner and the served encoder are the same method |
| Retrieval engine | Exact normalized inner product | Stable deterministic ranking across repeated queries |
| Warm query latency — CPU | p50 119.2 ms · p95 124.0 ms | 30-run within-language reference benchmark |
| Warm query latency — GPU | p50 25.7 ms · p95 26.9 ms | 30-run within-language reference benchmark |
| Score contract | Style · Topic/Tone · Affinity shown separately | A user can see which signal produced each part of the result |
| Text policy | Direct original-language comparison | Translation does not overwrite authorial syntax or register |
7 · Known limitations
- Similarity scores are uncalibrated cosines: valid for ranking, meaningless as probabilities or percentages.
- Cross-language matches are reduced-confidence and ranked per target language; raw multilingual cosines are not comparable across language pairs (Icard et al. 2025).
- Source-heldout evaluation is necessary but insufficient — author and topic can stay correlated across sources; cross-topic results are tracked separately.
- Central profiles can be over-returned for texts they did not write. Density-aware reranking reduced aggregate exposure concentration in exploratory cross-fitting but did not establish a ranking gain, so it is not deployed (Alipoormolabashi et al. 2025).
8 · References
- Kittler, Hatef, Duin & Matas (1998). On Combining Classifiers. IEEE TPAMI.
- Montague & Aslam (2001). Relevance Score Normalization for Metasearch. CIKM.
- Cameron, Gelbach & Miller (2008). Bootstrap-Based Improvements for Inference with Clustered Errors. Review of Economics and Statistics.
- Cawley & Talbot (2010). On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation. JMLR.
- Radovanović, Nanopoulos & Ivanović (2010). Hubs in Space: Popular Nearest Neighbors in High-Dimensional Data. JMLR.
- Schnitzer et al. (2012). Local and Global Scaling Reduce Hubs in Space. JMLR.
- Guo, Pleiss, Sun & Weinberger (2017). On Calibration of Modern Neural Networks. ICML.
- Geifman & El-Yaniv (2017). Selective Classification for Deep Neural Networks. NeurIPS.
- Dror et al. (2018). The Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing. ACL.
- Sagawa, Koh, Hashimoto & Liang (2020). Distributionally Robust Neural Networks for Group Shifts. ICLR.
- Sawatphol, Udomcharoenchaikit & Nutanong (2024). Addressing Topic Leakage in Cross-Topic Evaluation for Authorship Verification. TACL.
- Huertas-Tato et al. (2024). Isolating authorship from content with semantic embeddings and contrastive learning. arXiv:2411.18472.
- Terreau, Gourru & Velcin (2024). Capturing Style in Author and Document Representation. arXiv:2407.13358.
- Wang et al. (2024). Multilingual E5 Text Embeddings: A Technical Report. arXiv:2402.05672.
- Qiu et al. (2025). mStyleDistance: Multilingual Style Embeddings and their Evaluation. Findings of ACL.
- Kim, Zhang & Jurgens (2025). Leveraging Multilingual Training for Authorship Representation. EMNLP.
- Alshomary et al. (2025). Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer Layers. EMNLP.
- Alshomary et al. (2025). Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution. COLING.
- Alipoormolabashi, Patel & Balasubramanian (2025). Quantifying Misattribution Unfairness in Authorship Attribution. ACL.
- Icard et al. (2025). Embedding Style Beyond Topics: Analyzing Dispersion Effects Across Different Language Models. COLING.
- Man et al. (2026). Explainable Disentangled Representation Learning for Generalizable Authorship Attribution. ACL.
- Anand, Alshomary & McKeown (2026). iBERT: Interpretable Embeddings via Sense Decomposition. EACL.