Testing Jev as Android Tools for Codex
Jev and Gemini 3.5 Flash Lite answered quickly, but adding either to Codex’s Android workflow made tasks slower. Across two experiments totalling 90 runs, neither justified the extra decision layer.
My first look at Jev tested classification. I’d also tried Jev Ultrafast for browser control before moving on to Mobile Jev for Android. In my original repo review, Jev Ultrafast had 22,041 GitHub stars and Mobile Jev had 433.
The idea made sense to me: a classification model could score the available actions and choose one quickly. I wanted my main coding agent to do less work. If Codex could hand off a routine sequence, it wouldn’t have to read the screen and decide where to click at every step. That could save time and cost. I tested Jev, then a small LLM, with Codex checking their results.
What I tested
My existing workflow already uses Android’s accessibility data and batches predictable actions over USB, so Codex doesn’t need a screenshot for every decision.
Mobilerun Portal is an app installed on the phone. With Android’s accessibility permission, it gives the computer a structured list of screen controls, including their labels, positions and states. Its keyboard can put text directly into a field. For example, an agent can read a button’s label and location, or check what a text field contains.
In my test, Portal supplied screen information and text entry, while my existing tools still sent taps and swipes. Codex or the delegated model chose what to do next. I first tested whether Portal improved direct Codex control, then added Jev 1.13 to make the action decisions.
I used my own Pixel 3 XL running Android 12 over USB, with Portal’s cloud connection disabled. Model inference was hosted.
I used 18 cases adapted from AndroidWorld and MobileMiniWoB. Fifteen required actions: timers, Bluetooth settings, text entry, scrolling selection and forms with changing field order. Three were already satisfied, where the correct response was to do nothing. Each approach tried all 18, making 54 attempts in the first experiment. Starting states were reset, web fixtures seeded, and run order balanced.
The tables score the whole workflow, including any Codex rescue. Task time is the median for successful tasks requiring actions. It includes thinking, tool turnaround, phone reads, actions, recovery and an independent final check; setup and the three already-satisfied cases are excluded. Codex handled those three without delegation.
Jev needed rescuing and added time
| Codex’s tools | Passed | Task time |
|---|---|---|
| Existing tools | 18/18 | 34.3 s |
| Portal | 17/18* | 29.8 s |
| Portal + Jev | 18/18 | 51.6 s |
One Portal attempt was blocked before any task action.
Portal alone gave a modest improvement. On the 14 active tasks both approaches passed, it saved a median of 3.3 seconds, or 10.8%, against my existing tools.
Against Codex using Portal directly, adding Jev cost a median 21.4 seconds extra, or 79.1% more time, across 14 matched successes. These percentages compare each task with itself across approaches, then take the median; they aren’t ratios of the table’s medians.
Jev’s 18/18 required Codex to rescue three of the 15 active tasks. Jev claimed completion before reaching Bluetooth, tapped an offscreen item without selecting it, and repeatedly focused an empty form field. Codex finished those tasks. The perfect pass count came with extra work and a longer wait.
Given the interest around Jev, I was surprised by how little I wanted to use this setup. These were basic tasks, yet Codex still had to step in.
Flash Lite: even the successful runs were slower
I repeated the 18 cases with Gemini 3.5 Flash Lite and a fresh direct-Codex baseline, both using Portal. That made another 36 attempts, with the order alternating. This tests whether Flash Lite helped Codex in that run; it doesn’t rank Flash Lite against Jev from the earlier experiment.
| Codex’s tools | Passed | Task time |
|---|---|---|
| Portal | 18/18 | 33.3 s |
| Portal + Flash Lite | 17/18 | 48.3 s |
Every one of the 14 active tasks both versions passed was slower with Flash Lite. The median added time was 17.6 seconds, or 56.9%. The slowdown remained even with its failed run left out.
In the failed run, Flash Lite turned Bluetooth off and back on, then claimed completion. Codex corrected it, but the task required a single change: fixing the final state didn’t undo the extra toggles. Another run submitted a correct form three times before the loop detector stopped it. Codex checked the result and that task passed. No false Codex completion slipped through either experiment.
Why fast decisions didn’t save time
Jev’s median response was 319 ms; Flash Lite’s was 778 ms. Those are times for one decision, while a finished task needs phone observations, actions and verification too.
For one task, entering a name and pressing Submit, Codex sent focus, type and submit in one command. Flash Lite performed the same three actions through four model requests, with phone-state reads between decisions. Verified completion took 25.1 seconds directly and 45.4 seconds delegated.
Across the Flash Lite experiment, model requests totalled 70.3 seconds, while the executor’s 154 screen-state reads took 312.8 seconds. The executor refreshed state to avoid acting on stale information. In this implementation, those repeated phone checks cost much more time than inference, while Codex could already batch predictable actions.
Flash Lite took over almost all the tapping and typing, but Codex still had to start tasks and check results. The recorded Codex controller commands fell only from 55 to 52 across the 18 cases. Those are tool commands, not model turns or tokens, so they don’t tell us how much Codex cost was saved.
The executor calls themselves were cheap: about $0.00525 for Jev’s 57 requests and $0.056 for Flash Lite’s 84. Without per-task Codex usage, I can’t show that the combined workflow was cheaper.
Would I adopt it?
Before testing, I set a minimum adoption threshold of 25% faster completion and five seconds saved, without an observed reliability loss. Neither model met it. Direct control remains the default. Portal’s two completed form comparisons saved about 13 seconds each, so it is worth keeping available for form input.
These are small workflow tests on one phone, not official AndroidWorld scores. The same Codex conversation learned across runs, briefs changed, and observation formats differed. The original baseline also had avoidable screenshot and typing delays.
Applying the action batching I already use with Codex might make Jev a useful improvement. I’d be willing to spend time on that if it could halve task time or make the workflow ten times faster. But setting up and tuning extra tools for a possible marginal gain isn’t worth the time to me. These Android tests didn’t show a speedup, and I’m not spending more time trying to use Jev for this workflow, at least for now.