Wissen's Retrieval System
- Headline statistics to showcase Wissen performance relative to others
- E.g. "On a sample of 2k single equity multi-year financial research questions, Wissen's true recall (@top 3-k) ranks number 1 vs alternatives (93% vs competitor ave of 68%). Within a 3-step workflow this places Wissen at 80% mean recall vs competitive ave mean recall 31%."
- Headline statistics to showcase Wissen's performance relative to our own history
- Key areas we are hill climbing on for the future
Evaluation Methodology
The Wissen Way
- Insert the usual methodology for retrieval evaluation.
- Insert our methodology for retrieval evaluation.
- Insert why our methodology differs to the usual, and what that means in the context of our domain and why it is important
Internal Benchmark Performance
Where we rank vs our own history, and what the future looks like for Wissen, Retrieval, and it's application in the Financial Sector
- Insert our latest headline performance on the key evaluation metrics that we care about
- Insert a section comparing our performance on those key evaluation metrics 1 month (or a set time horizon) ago
- Show areas where we still underperform but are actively conducting effective R&D to hill climb on and improve
Relative Benchmark Performance
Wissen vs Alternatives - Why is what we're doing differentiated and important?
- Rank Wissen Retrieval vs credible alternatives (such as llama Index, Azure, Vertex AI, AWS, OpenAI, Weaviate) using Wissen's evaluation criteria. This is the first order mark to reality of our retrieval efforts on a use-case by use-case manner.
- Rank Wissen vs Direct Competitors (such as AlphaSense, Bloomberg, FinChat, Quartr, FinTool, Hebbia, Rogo, Balyasny, DESCO, Marshall Wace). This will require creativity and gumption, but fundamentally is what we should be itching to understand. All those hours we are putting in, are we on track to provide surplus value relative to our direct competition? If so where and how?Who are the rent extractors in this space and should our value add be elsewhere?
- Rank Wissen vs soft comps (such as ChatGPT, Perplexity, Grok, Notebook LM etc.). This is not the cleanest comparison, but a valuable one for a large segment of the market who are trying to understand really simply are we better than just getting a ChatGPT Pro or Perplexity Pro.