Jev is overhyped as a solo tool, but solo is not the point
- But I think what’s being missed is this: Jev is going to be one of a number of instances in the near future where top level LLMs sub out to more specialized tools. Jev’s error rate as an adjudicator is too high for a lot of tasks, and it does not supply evidence or citation. But as a screen it can take the burden off the LLM and route to the LLM + search the issues that require true investigation. That radically shifts what is financially possible. To some extent there is hype around Jev because the ability to put it in a frontier LLM harness as part of a workflow addresses a lot of the problems that old classifiers had.
This is worth thinking about for spam/misinfo detection
Tasks
- 3:00 – FA
- 5:30 Espaniol
- Put together a list of venues and see what they want submitted
SBIRs
- 9:00 Standup
- More WG development. I’d really like to get interaction with an OpenRouter model
- Should probably look at storing chat histories in the project?
- Put the email chain with Matt in the Overleaf
- Embeddings today?
