Capital Allocations at Wissen
The post here is intended to serve as a reply to Moment to drive momentum focussed on two things
- Actionable Edits - it will provide what + why I'd edit the plan on what work needs to be done by who and why alongside with any additions/breakdowns
- What R&D actually means (and has meant) for us - how I think we should be thinking about the way we 'spend' time and have 'spent' time especially on the topic of 'R&D'
- As we move on from the semi-sporadic R&D... - this is more of a 'next steps' + more concrete answers on next steps and 'things to be done'
Actionable edits
Day 1-3 (72h)
- (Correction) Draft approach to evaluations in parsing & search and functionally what changes mean [Owner: King]
- Why: I've just been a lot more in the weeds of the components so it's going be easier if I release+structure certain components that needs to be eval'd. The distinction I'd draw is:
- set-up - setting up the structure of evals should be on me (particularly as I'm extremely pedantic about the structure of it, see more context at bottom) and you can also see in 'Drafts' a taster for what I mean in the 'what do we evaluate' document'.
- domain oversight - as the structure is created so that it's easy to append+add metrics based on a traceable+reproducible source truth, that is where your value can be maximised which is where Owner: King -> Hamza
- note: there may be two different 'types' of evals i.e. more on the deeper LM post-retrieval processing and then retrieval+parsing evals (both are sort of necessary so the above may be a given already)
- Why: I've just been a lot more in the weeds of the components so it's going be easier if I release+structure certain components that needs to be eval'd. The distinction I'd draw is:
- (Agree) Write-up why we've spent time on retrieval, what has been done, and what it unlocks downstream [Owner: King]
- (Agree) Draft initial set-up of the retrieval [Owner: King]
- (Agree) Run 'initial evaluations' on retrieval to validate time spent [Owner: King]
- (Change) Product Specification & initial demo build outs [Owner: Hamza]
- Draft up a spec' on what is the hypothesise on the functional utility of a 'Wissen Agentic QA' system including [Owner: Hamza]:
- the definition of an 'agentic' + 'qa' system
- examples of use-cases in that system
- the functional requirements to satisfy the utility
- what does 'initial version' mean and what does an 'evolved version' look like so we can be aware that we are 'sizing' into this trade
- what does the 'marketing' of this product look like to customers once it's completed i.e. how do we interface it and/or different ways we could interface it (via human, via interface) and the current skew we should have
- why: This will give a strong blueprint around the 'why' behind the build-out and the ability to 'close-the-loop' as soon as the product is rolled out which will need a lot of intentionality + thinking.
- strong preference: spec'ing out the details of these features will be the deep thinking + hard work that needs to be done to avoid tangents and is best if you take command over it (especially the latter parts of 'how it's presented, what is shown, what is mentioned, and the general variants). Once those spec's are there, the 'implementation' should be a lot more trivial as it'll give me room to be 'creative' in implementation to reduce time to reach the same functional utility we desire. I won't be as pedantic on what (value prop) rather the how (method) will always have strong opinions - both requires deep thinking and the former you will naturally maximise your zones and latter you will be able to leverage my zones.
- Find inspiration (I see you've started ;) ) & build initial drafts of the demo version [Owner: Hamza]
- Comment: we need to figure out how to minimise integration costs as taking it to production + wrap-up will likely fall on my lap. I'll let you pull me in as you need it so I can focus on closing the loops in next 3d from my end
Day 4
- (Agreed) let me know the 'best' means to communicate reminders as I tend to take a hand-off approach to give flexibility
Day 5-7
- General: my work so far has not been on 'real R&D' rather been trying to get to functional versions to generate some internal conviction that our approach works on non-negotiable requirements as such will continue work on that front, more context below.
- (Question 1) If you had to think about what we unlock at a basic Q&A stage -> a more evolved complex (Q&A) i.e. agentic stage/report writing, what gets unlocked at what stage and why are earlier stages not unlocking value yet? This is primarily to flesh out a bit more whether we are potentially adding more requirements than necessary to start shipping versions that actually already works and we have
- (Question 2) What would you like from me on the UI + Presentation side before you send off things to design partners? Or would you purely be sending locally generated outputs to design partners (i.e. like earnings views pre-print)?
Day 8
- (Agree + Strong Suggestion) Consultancy - yes, please do also think through under a 'consultancy' route, what does 'improving' tech fundamentally improve for us (i.e. is it realistically time to answer questions) and also the conjectured RV we can capture as a result of it. This does not (and should not) need to be a 5y horizon bet, rather it's to generate some functional short-term alignment (next months) on what tech improvements means internally e.g. overtime it's <hrs spend per customer whilst top line remaining the same/growing (as such this is the initial internal KPI we track). This also impacts the UI/infra side because if we take Michel's approach i.e. they ask us (you) for things via email -> we push unidirectionally to a dashboard, it will massively simplify our requirements by a lot whilst also giving us flexibility on not needing to spend as much time on building strong infra in the short-term and size into that infra as well
What R&D actually means (and has meant) for us
- R&D from my perspective (measured by time) has meant
- 1) setting up a functional version(s) to test;
- (I've been massively able to speed this up due to working independently so can take much more 'hacky' methods to get to versions for you)
- 2) evaluating that functional version;
- Most of the time has actually been on (1) rather than heavily (+being rigourous) on (2)
- 1) setting up a functional version(s) to test;
- (1) is important + takes time because we have implicit/explicit non-negotiables e.g. we can search over diagrams on company presentations (alongside other documents), we can parse company presentations and that text is given to an LM, data is complete (no missing 10-Qs or transcript in a fiscal period), fiscal years + quarters + fiscal dates are in place where it is necessary (10-Ks/10-Qs etc.) and not missing and accurate, it doesn't break >20% of the time when using it and more. These all gets baked into the 'overarching initial retrieval requirements'. These non negotiatiables are part of what we incur before we actually evaluate with rigour (latter is where the 'questionable' abstract R&D starts coming into thought that 'blocks real business progress', former is what we need anyway to have 'something' to show). These non-negotiables are also a learning on my end as we've been cooking.
- (1) and (2) in some sense was to answer the question of
- is it even technically feasible to be able to reliably search over company presentations and all the other documents types at the point of a query and had off to an LM? The assumption we had to this was always a strong yes.
- Then the real (implicit) question was: do we need a team of 40> to build this out (i.e. does it make sense to do it now)? Which is where the 'R&D hypothesise' which was formalised came to exist i.e. 'what if we indexed all the information based on images'.
As we move on from the semi-sporadic R&D...
- As we've garnered greater conviction on a variant of (1), there are certain things that don't change and needs to change
- (don't change) non negotiables
- (don't change) functional version that doesn't break constantly
- (change) time to a '5-10x better' version shouldn't mean overhauls especially if the premise of the system is the same; time to onboarding new equities should not incur XX hours of my time; time to knowing 'why' something is wrong should be knowable (to the extend we can know) so that fixes shouldn't take XX hours rather minutes; (and more)
- (overall change) the overall change in this process has been to
- reduce the variable costs; reduce the variable costs that produce more variable costs (patch works can cause more bugs)
- (change) as enter into a phase where we see very strong signs of life, and will end up hill-climbing on outputs, it's important to ensure that we approach the evaluation set-up (not the actual domain evaluations nor hill-climbing itself, rather the set-up) with a lot of rigour and build infra around it. This means 'basic' things like ability to reproduce results, ability to trace where a evaluation metric came from (especially if it's an average of results), extending the test sets, ease of compatibility with other evaluations (non-LM) to make the interface much more manageable, etc.
- (agreed change prompted by you) start gearing up and steel-manning what is it actually that is being shipped and shipping it ahead of Malaysia with enough time to QA the system to build conviction on it's abilities
- (agreed) R&D is 'bound to take time' always as such we can leverage the async environment we get over mid-may to start of june to distribute work
- My time allocation over next 8-14 days including what you mentioned (no specific days rather just time costs but will prio ones that unblock you + alrdy near completion)
- automate most of the equity onboarding process (95% there) - <3-4h left
- automate checking data integrity/completeness that devs send (80%) - <3h left
- structure the set-up of experiments with write-ups (60%) - <2d
- run initial experiments and write-up results (10%) - <1d
- onboard an industry e.g. consumer staples worth of equities - <2d (less variable in time but rather an allocation incase bugs happens)
- wrap this into an 'initial version' of the product <2d (depending on how well the product is spec'd out + how many requirements we're placing in there)
- note: the things that will cause delays is if there are unknown unknown bugs and any known known bugs, I'm automating away