Under the Hood, part 1 · October 2026 · Steepwind Labs engineering
Panda gets reflexes
We put a fast classifier in front of the LLM in our phone agent. Most screen decisions now take about 0.4 seconds, and Panda knows when to hand the decision to something smarter.
Key numbers
- 0.38 s median Jev decision per screen step (p90 0.55 s).
- 97.6% of confident step picks were correct.
- 85% of steps were confident enough to skip the LLM.
- 80 of 80 confident voice routes were correct, including Hinglish and speech-recognition typos.
What Jev is
Jev is a "System One" model from TypeSafe. It does not generate text. Given a state and a fixed set of typed questions, it answers all of them in one pass with a probability for each option. Panda already knows every legal move on the screen, so Jev only has to pick one.
Where it sits in the loop
Each step, Panda reads the screen, turns it into a short list of candidate moves, and asks Jev to choose. Our code decides whether to trust the answer. If Jev is sure, Panda acts right away. If not, the step goes to the LLM. Jev never writes text, never ends a task on its own, and only counts as sure when its first choice clearly leads its second.
What testing showed
We replayed 130 steps from 49 recorded runs on an Android test phone (98 labelled). When Jev's top two options were close it was right one time in six; with a clear lead it was right 97–100% of the time. Its "goal done" signal overlapped too much between finished and unfinished steps to end runs by itself. On 28 app-pick prompts across 165 installed apps it was right on all 26 it was decisive about. For voice requests it routed 85 of 87 correctly.
What we learned
Use a classifier where the options are known. Calibration matters more than raw accuracy. Record every run from day one. Once decisions are fast, the remaining latency is in connections, warm-up and screen animations.