5 Commits

4 changed files with 223 additions and 166 deletions

View File

@@ -1,12 +1,53 @@
# kongruity
“...All those moments will be lost in time, like tears in rain.”
kongruity pulls in unstructured artifacts of the creative-engineering process -- to-dos, action items, agile tickets, Jira thread comments, Slack thread comments, retrospective notes -- and synthesizes them into semantically coherent, prioritized clusters that can be incorporated into implementation planning.
In kongruity world, the artifacts become "sticky notes." A board full of them looks chaotic. With a click, the LLM analyzes their semantic meaning and groups them into thematic clusters, each with a descriptive header.
In kongruity, the artifacts become "sticky notes." A board full of them looks chaotic.
## Result evalutation/validation:
With a click, they are semantically evaluated, grouped into thematic clusters with descriptive headers, rankable and exportable to project planning and execution tools.
An independent RAG embedding-based model scores evaluates quality using cosine similiarty, so the output is qualitatively refined. From there, teams can drag-and-rank related task clusters by implementation priority - turning noise into an actionable workflow.
## Clustering and evaluation: methodology
Two models run in parallel, and neither sees the other's work. Anthropic's `claude-sonnet-5` (`backend/services/clustering.service.js`) reads the raw text of every note and groups them into labeled thematic clusters.
At the same time, Voyage AI's voyage-3 model (`backend/services/embedding.service.js`) converts each note's text into a numeric representation of its semantic meaning aka vector.
Once the LLM returns, kongruity scores that grouping (`backend/services/validation.service.js`) using an established silhouette coefficient, with cosine distance rather than Euclidean as the proximity metric.
For each note, it weighs the average distance to the other notes in its own cluster against the average distance to the notes in the nearest neighboring cluster. Averaged across every note, this yields a single cohesion score in the range [−1, 1], displayed at the top of the results.
This yields an empirical groundedness evaluation. One model proposes the grouping; an independent model evaluates grouping accuracy.
Note that: before scoring, structural validation confirms that each note landed in exactly one cluster, that no cluster is empty, and that no hallucinated note IDs appear. A malformed response to the validation completely fails, rather than quietly returning a partial board.
## Reading the cohesion score
Average silhouette width is a widely-used measure of clustering quality. Higher values indicate:
1. The qualitative semantic cohesiveness of clusters, and:
2. How well-separated each cluster is from its nearest neighboring cluster.
Although the coefficient is mathematically bounded by [−1, 1], cosine distance between high-dimensional text embeddings is compressed: unrelated notes sit close to orthogonal, so both the within-cluster and nearest-cluster distances land near 0.8. Because silhouette divides the gap between them by the larger of the two, the practical range on embedding data is roughly [−0.05, 0.10] rather than the full interval.
The bands below are therefore calibrated against that observed range. On the seed board, the five ideal thematic clusters score 0.09; swapping a few notes between clusters drops it to 0.06; a scrambled assignment falls below zero.
The score appears above the results with a plain-language band:
- **0.07 and above** — Strong
- **0.04 to 0.06** — Moderate
- **0.01 to 0.03** — Weak
- **Below 0.01** — Poor
A score near 0.00 means the grouping is no better than chance. Bands are specific to `voyage-3` cosine distance and would need recalibration behind a different embedding model. See Hugo Sträng, Tai Dinh. An upper bound on the silhouette evaluation metric for clustering. Pattern Recognition, Volume 178, 2026, 113402, ISSN 0031-3203.
## Organizing clusters, exporting to workflow software
Teams can drag-and-rank related task clusters by implementation priority - turning noise into an actionable workflow.
(Integrations with third-party project management, planning and workflow applications are action-items for next major version, see Roadmap, below)
## How it works
@@ -18,12 +59,30 @@ An independent RAG embedding-based model scores evaluates quality using cosine s
## Dev implementation notes
As of 03.04.2026, two of the above-described features are in the planning and implementation phase:
Developers may swap in other LLM SDKs/APIs and alter prompt syntax in `backend/services/clustering.service.js` to experiment with LLMs and platforms of their choice.
1. **Ingestion** - The implementation goal is a system for easily tagging items/issues mentioned in Slack, Jira comments, etc. (similar to hashtagging), and running batch "pulls" of these tagged items into kongruity via REST API interface.
2. **Prioritization** -- The final prioritization and planning stage will ultimately result in "pushing" these ordered items back **out** into Jira, Asana, Rally (etc.), for incorporation into Epic/Sprint workflows.
## Development Roadmap
Also, developers may swap in other LLM SDKs/APIs and alter prompt syntax in `backend/services/clustering.service.js` to experiment with any model or platform of their choice.
### Ingestion — pulling tagged artifacts in
- [ ] **Slack** — where decisions actually get made; a `:sticky:` emoji reaction fires an Events API webhook that pulls the message in.
- [ ] **Microsoft Teams** — same capture gesture for enterprise shops; message extension plus Graph change notifications.
- [ ] **Jira** — label- or mention-triggered webhook scoped by JQL. (This is where comments typically carry half the backlog's context.)
- [ ] **Linear** — engineering-side tickets and threads; label-triggered GraphQL webhook.
- [ ] **GitHub** — issue, PR review, and discussion comments; label- or mention-triggered webhook.
- [ ] **Miro / FigJam** — REST API import
- [ ] **Confluence / Notion** — page and inline-comment fetch. (Where retro and planning notes are born).
- [ ] **Meeting transcripts (Granola, Otter, Zoom, Google Meet)** — where retros are now recorded, an option for action-item extraction from the transcript API.
- [ ] **Generic REST, email, and Zapier** — authenticated bulk `POST /v1/notes`.
### Export — pushing ranked clusters to workflow tools
- [ ] **Jira** — drag-rank written through the Agile API's board rank endpoint.
- [ ] **Asana** — drag-rank written as task order within the section.
- [ ] **Rally** — drag-rank written as portfolio rank. (Clusters become features and notes, which become stories).
- [ ] **Linear** — drag-rank written to issue `sortOrder`.
- [ ] **Azure DevOps / GitHub Projects v2** — drag-rank written as project field ordering. Clusters become work-item parents.
- [ ] **CSV, JSON, and Markdown** — direct download from the cluster view. (Should ship before any OAuth work.)
## Prerequisites
@@ -157,5 +216,3 @@ This runs Vitest with jsdom. For watch mode during development:
```bash
npm run test:watch
```

View File

@@ -1,8 +1,7 @@
import { Transform } from 'node:stream';
/**
* Groups an object-mode stream into fixed-size arrays. Backpressure on the
* readable side is what limits how many batches are ever in flight.
* Backpressure on readable side is limits how many batches are in flight.
*
* @param {number} size - maximum items per emitted batch
* @returns {Transform}
@@ -42,9 +41,6 @@ export const batch = (size) => {
};
/**
* Serializes an object-mode stream into a JSON array, one element at a time,
* so no complete copy of the payload is ever held in memory.
*
* @returns {Transform}
*/
export const jsonArray = () => {

View File

@@ -6,10 +6,14 @@ import Sticky from './sticky';
import Button from './button';
import '../styles/stickies.css';
// In practice, silhouette on cosine distance between text embeddings occupies roughly
// [-0.05, 0.10], not strict theoretical [-1, 1]: near-orthogonal vectors put both the within- and
// nearest-cluster distances close to 0.8, and the coefficient divides their gap
// by the larger. These bands are calibrated to that range for voyage-3. see README, Reading the cohesion score
const scoreLabel = (score: number): string => {
if (score >= 0.7) return 'Strong';
if (score >= 0.4) return 'Moderate';
if (score >= 0.1) return 'Weak';
if (score >= 0.07) return 'Strong';
if (score >= 0.04) return 'Moderate';
if (score >= 0.01) return 'Weak';
return 'Poor';
};

View File

@@ -14,7 +14,7 @@ const MOCK_CLUSTER_RESPONSE = {
{ label: 'Auth Issues', noteIds: ['note_001'] },
{ label: 'Export Issues', noteIds: ['note_002'] },
],
score: 0.74,
score: 0.09,
};
let fetchMock: ReturnType<typeof vi.fn>;