Re-aligned heuristic, updated readme, added code comment

This commit is contained in:
KS Jannette
2026-08-01 04:06:39 -04:00
parent ac75a30b61
commit 6f09b6ecdc
3 changed files with 119 additions and 111 deletions

View File

@@ -30,14 +30,18 @@ Average silhouette width is a widely-used measure of clustering quality. Higher
2. How well-separated each cluster is from its nearest neighboring cluster.
Although the coefficient is mathematically bounded by [−1, 1], cosine distance between high-dimensional text embeddings is compressed: unrelated notes sit close to orthogonal, so both the within-cluster and nearest-cluster distances land near 0.8. Because silhouette divides the gap between them by the larger of the two, the practical range on embedding data is roughly [−0.05, 0.10] rather than the full interval.
The bands below are therefore calibrated against that observed range. On the seed board, the five ideal thematic clusters score 0.09; swapping a few notes between clusters drops it to 0.06; a scrambled assignment falls below zero.
The score appears above the results with a plain-language band:
- **0.70 and above** — Strong
- **0.40 to 0.69** — Moderate
- **0.10 to 0.39** — Weak
- **Below 0.10** — Poor
- **0.07 and above** — Strong
- **0.04 to 0.06** — Moderate
- **0.01 to 0.03** — Weak
- **Below 0.01** — Poor
Silhouette values are archetypically bounded below 1.0 for real-world data, so the number is best read as a relative measure. See Hugo Sträng, Tai Dinh. An upper bound on the silhouette evaluation metric for clustering. Pattern Recognition, Volume 178, 2026, 113402, ISSN 0031-3203.
A score near 0.00 means the grouping is no better than chance. Bands are specific to `voyage-3` cosine distance and would need recalibration behind a different embedding model. See Hugo Sträng, Tai Dinh. An upper bound on the silhouette evaluation metric for clustering. Pattern Recognition, Volume 178, 2026, 113402, ISSN 0031-3203.
## Organizing clusters, exporting to workflow software