Phil 9.22.2026

Jev is overhyped as a solo tool, but solo is not the point

  • But I think what’s being missed is this: Jev is going to be one of a number of instances in the near future where top level LLMs sub out to more specialized tools. Jev’s error rate as an adjudicator is too high for a lot of tasks, and it does not supply evidence or citation. But as a screen it can take the burden off the LLM and route to the LLM + search the issues that require true investigation. That radically shifts what is financially possible. To some extent there is hype around Jev because the ability to put it in a frontier LLM harness as part of a workflow addresses a lot of the problems that old classifiers had.

This is worth thinking about for spam/misinfo detection

Tasks

  • 3:00 – FA
  • 5:30 Espaniol
  • Put together a list of venues and see what they want submitted

SBIRs

  • 9:00 Standup
  • More WG development. I’d really like to get interaction with an OpenRouter model
    • Should probably look at storing chat histories in the project?
  • Put the email chain with Matt in the Overleaf
  • Embeddings today?