Compare commits
14 Commits
FEAT-user-
...
9708d36d51
| Author | SHA1 | Date | |
|---|---|---|---|
| 9708d36d51 | |||
| 247ba97a3a | |||
| fc6ceb558b | |||
| 676c86cf43 | |||
| 6810704173 | |||
| 4a1c67c518 | |||
| 3a4ac641cd | |||
| 861983531a | |||
| fe36ea6aaa | |||
|
|
62dd6de60c | ||
| d481b79465 | |||
| b8c65348e8 | |||
| d6902daa49 | |||
| a79c5ad1cc |
69
README.md
69
README.md
@@ -1,99 +1,156 @@
|
||||
# Citation Sentinel
|
||||
# Citation Sentinel - 2025 @sjDev - LICENSE: MIT
|
||||
|
||||

|
||||
|
||||
A source-grounded research assistant that empowers users to transform an otherwise-unmanageably-large corpus of data on any topic of interest into refined, easily-digested subtopics and pose focused inquiries, to instantly receive concise, quantifiably-evaluated, accurate responses. The system also provides suggested follow-up questions further refining in user inquiries by detecting the query goals.
|
||||

|
||||
|
||||
Image one: source-grounded research assistant for applied use of LLMs.
|
||||
|
||||
**Citation Sentinel leverages multiple Large Language Models for analyzing and conveniently working with large data corpora, when query-result accuracy is of the highest value. Examples include complex litigation (i.e. Patent/IP), technical medical data exploration (reviewing health records, test results, or studies to find hidden patterns, trends, and answers), advanced LLM development and refinement.**
|
||||
|
||||
Citation Sentinel empowers users to transform an otherwise-unmanageably-large corpus of data on any topic of interest into refined, easily-digested subtopics and pose focused inquiries, to instantly receive concise, quantifiably-evaluated, accurate responses. The system also provides suggested follow-up questions further refining in user inquiries by detecting the query goals.
|
||||
|
||||
ALSO NOTE IN DEMO: clickable inline citations (in blue) that take user to source-grounded research basis/knowledge corpus supporting query response, with rated “groundedness” score.
|
||||
|
||||
|
||||
# Reliability and Understandability
|
||||
|
||||
|
||||
Citation Sentinel's Retrieval-Augmented Generation (RAG) pipeline uses two-stage pass-through to Voyage AI embedding and ranking models, which uses cosine similarity methodology to scoring to evaluate query-source embeddings in multi-dimensional vector space. This yields a "groundedness" score.
|
||||
|
||||
|
||||
Groundedness metrics quantify how well an AI-generated answer is supported by retrieved context. It measures "faithfulness" to source documents, ensuring the answer is not hallucinated, inaccurate or stale (pulled from the model's training data).
|
||||
|
||||
|
||||
# Architecture - Basic Overview
|
||||
|
||||
Built with a React/Vite frontend and a Typescript/Node/Express backend, using Anthropic Claude for generation, OpenAI Whisper for video audio track transcription, Voyage AI voyage-3 for embeddings and Voyage AI rerank-2 for response cosine similarity scoring, to generate the "groundedness" score, presented as an easily-understandable red/yellow/green "badge" style tooltip in the response (see above.)
|
||||
|
||||
Built with a React/Vite frontend and a Typescript/Node/Express backend, using Anthropic Claude for NL response generation, OpenAI Whisper for video audio track transcription, Voyage AI voyage-3 for embeddings and Voyage AI rerank-2 for response cosine similarity scoring (to generate the "groundedness" score) presented as an easily-understandable red/yellow/green "badge" style tooltip in the response (see image above.)
|
||||
|
||||
|
||||
## Query pipeline
|
||||
|
||||
|
||||
Source Ingestion -> Parsing -> Chunking -> Embedding (using Voyage AI voyage-3 model) -> Storage (vector store) -> { user query submission } -> Evaluation of User Query -> Retrieval -> Ranking -> Response Generation (Using Anthopic's claude-opus-4-6 model) -> Response Groundedness Scoring (using Voyage AI rerank-r model)
|
||||
|
||||
## Why Database Math Can't Replace Groundedness Scoring
|
||||
|
||||
Many recent additions to the Vector database/store product space use distance metrics (Cosine, L2, Inner Product) for Bi-Encoder Similarity. They compare the vector of the user query against the vectors of chunks independently.
|
||||
|
||||
The rerank-r model at the end of the Source Sentinel pipeline performs Cross-Encoder Evaluation, taking two entirely separate text inputs:
|
||||
|
||||
1) The generated response from claude-opus-4-6 and
|
||||
|
||||
2) The raw source chunks
|
||||
|
||||
It processes them simultaneously through deep attention layers to check for hallucinations, missing context, and factual alignment.
|
||||
|
||||
A vector database cosine metric only calculates how close two embeddings are in coordinate space. It **does not read the generated text of a response to verify if it accurately reflects the source chunks**.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
|
||||
- **Node.js** (v18+)
|
||||
- **yt-dlp** -- Required for YouTube video source support (`brew install yt-dlp`)
|
||||
|
||||
|
||||
## Getting Started
|
||||
|
||||
|
||||
```bash
|
||||
# clone and install
|
||||
git clone <repo-url> && cd citation_sentinel
|
||||
cd server && npm install && cd ..
|
||||
cd client && npm install && cd ..
|
||||
|
||||
|
||||
# configure
|
||||
cp server/.env.example server/.env
|
||||
# edit server/.env and add your ANTHROPIC_API_KEY, VOYAGE_API_KEY,
|
||||
# and optionally OPENAI_API_KEY (only needed for audio source transcription via Whisper)
|
||||
|
||||
|
||||
# run
|
||||
make dev
|
||||
```
|
||||
|
||||
|
||||
## Rationale
|
||||
|
||||
|
||||
This app demonstrates the core source-grounded Q&A pattern with transparent retrieval, generation, and groundedness scoring (LLM response quality cosine scoring) -- all with swappable models and is fully open source.
|
||||
|
||||
|
||||
## Methodology
|
||||
|
||||
|
||||
This is a RAG (Retrieval-Augmented Generation) and query-response source-groundedness assurance system allowing users to create a research knowledge corpus, including:
|
||||
|
||||
|
||||
1. Documents — PDFs, DOCX files, plain text audio files.
|
||||
2. Internet sources: via web URLs.
|
||||
2. Audio sources: i.e. YouTube videos, audio from which is transcribed to text.
|
||||
|
||||
|
||||
These are then parsed, split into ~2000-character overlapping chunks, and converted into vector embeddings using Voyage AI's voyage-3 model.
|
||||
|
||||
|
||||
The embeddings are stored in an in-memory vector store.
|
||||
|
||||
|
||||
When a user submits a query, the system enforces groundedness through a multi-layered strategy:
|
||||
|
||||
|
||||
1. **Retrieval constraint** — The query is embedded (via Voyage AI voyage-3) and compared against stored chunk vectors using cosine similarity, returning the top 20 candidates.
|
||||
|
||||
|
||||
2. **Reranking for precision** — Those 20 candidates are sent to Voyage AI's rerank-2 cross-encoder, which re-scores each query-chunk pair with deeper semantic analysis. Only the top 5 survive.
|
||||
|
||||
|
||||
3. **Prompt-level constraint** — The top 5 chunks are passed to the "Primary LLM" (Claude opus-4-6) for natural language query response. The Primary LLM must cite sources using bracketed indices (e.g., [1], [2]) and admit when sources are insufficient.
|
||||
|
||||
|
||||
4. **Schema enforcement** — The Primary LLM's response is constrained to a JSON schema requiring structured fields (answer, citedSourceIndices, followUpQuestions). Any cited source indices that do not correspond to real source groups are programmatically removed.
|
||||
|
||||
|
||||
5. **Post-generation groundedness scoring** — The answer is split into individual sentences. Voyage AI voyage-3 embeds each sentence, and compares it (using cosine similarity) against the vectors of the cited chunks. The similarity is calibrated to a 0–1 scale and averaged, producing a groundedness score surfaced to the user with a visual indicator (“badge” in green/yellow/red).
|
||||
|
||||
|
||||
## Similarity Metrics for Semantic Understanding
|
||||
|
||||
|
||||
Cosine similarity measures how closely two vectors (representing data like words, images, or preferences) are aligned in a multi-dimensional space by calculating the cosine of the angle between them.
|
||||
|
||||
|
||||
Archetypical cosine scores range from -1 to 1.
|
||||
|
||||
|
||||
1: Vectors point in the exact same direction (highly similar).
|
||||
0: Vectors are at a 90-degree angle (orthogonal/unrelated).
|
||||
-1: Vectors point in opposite directions.
|
||||
|
||||
|
||||
Cosine distance between high-dimensional text embeddings is compressed: unrelated notes sit close to orthogonal, so both the within-cluster and nearest-cluster distances land near 0.8. In this application, scores are normalized to a bounded range wherein unrelated query-response pair scores are nearly orthogonal.
|
||||
|
||||
|
||||
## Design
|
||||
|
||||
|
||||
Two-package monorepo:
|
||||
|
||||
|
||||
- `server/` -- single Express backend with layered architecture (routes -> services -> stores). Routes orchestrate; services contain business logic; stores manage in-memory state.
|
||||
|
||||
|
||||
- `client/` -- React 19 SPA via Vite. Two-panel layout: sidebar for data sources, main area for LLM chat with explorable citations, groundedness badges (cosine similarity scoring of LLM responses), and follow-up question chips.
|
||||
|
||||
|
||||
## License
|
||||
|
||||
|
||||
MIT. See [LICENSE](./LICENSE).
|
||||
|
||||
|
||||
## Author
|
||||
|
||||
|
||||
@sjdev
|
||||
|
||||
|
||||
|
||||
@@ -104,11 +104,11 @@ export async function generateStudyGuide(sourceGroups: SourceGroup[]): Promise<S
|
||||
const sources = buildSourceBlock(sourceGroups);
|
||||
const start = Date.now();
|
||||
|
||||
const prompt = `You are an expert educator. Given the source documents below, produce a comprehensive study guide. Follow these rules:
|
||||
const prompt = `You are an expert reasearch assistant. Given the source documents below, produce a comprehensive guide. Follow these rules:
|
||||
|
||||
1. Create a structured outline of the main ideas organized into logical sections.
|
||||
2. For each section, provide: bullet-point key concepts, key terms with definitions, and 2-3 self-review questions.
|
||||
3. At the end, provide mnemonic devices or simplified restatements to help memorize the hardest concepts.
|
||||
3. At the end, provide mnemonic devices or simplified restatements of the hardest concepts.
|
||||
4. Be concise but thorough. Use simple, clear language.
|
||||
|
||||
--- SOURCE DOCUMENTS ---
|
||||
|
||||
Reference in New Issue
Block a user