BENCHMARKS
Benchmarks.
Measured results for computer use and browser use on real, owned hardware: how long tasks take, where clicks land, how fast control changes hands, and how much memory ibara uses. Each number has its date and machines, and the method says how the full benchmarks are run.
Updated
Measured so far#
These numbers come from runs on the tested hardware, tater0 and tater1: five-year-old mini PCs with 10-watt Celerons. Both ran Omarchy 4.0.4 with Hyprland 0.56.2; Cua’s driver was 0.29.1 on tater0 and 0.28.2 on tater1. On 28 September real agents ran the first full benchmark, and the handoff, input and memory measurements were repeated. On 29 September the same tasks ran again on the published release, 0.1.0-17. The earlier acceptance runs stay below, dated, and the acceptance runs are also listed, passed or not, in the test results.
Tasks on the release#
On 29 September, from 02:56 to 06:32, real agents ran the same 12 tasks and checks as on 28 September on the published package, ibara 0.1.0-17, the same on both machines. tater0 had Google Chrome and a wireless mouse receiver plugged in; tater1 had Chromium, no mouse, and was freshly installed from GitHub that night. Each agent had only the ibara MCP server, the packaged ibara skill and the ibara note in its instructions; the project’s own source and notes were hidden from it. Every run counts.
- 132 of 157 runs passed (84 %), against 49 of 73 (67 %) on 28 September. Desktop tasks 56 of 74, browser tasks 76 of 83; 18 runs hit the time limit. tater0 67 of 81, tater1 65 of 76.
- Failures: 13 were ibara’s, 8 were the agents’, and 4 were the benchmark’s own setup, fixed during the run and still counted. The weakest task, the file manager (3 of 13), failed because the rename popup took the keyboard and refused every input; that caused 10 of the 13 ibara failures. Some clicks on its buttons were also refused along the way, which cost the agents extra calls. The other 3 ibara failures came from the editor keeping its settings where the agent’s usual tools can’t see them; since 0.1.0-20 the editor uses your own Mousepad settings. ibara also sometimes marked finished browser tasks incomplete, because its text check read only part of the page. 0.1.0-18, released the same morning, fixes the popup, the refused clicks and the text check; a focused rerun of 5 tasks on it is below, and the full benchmark hasn’t been run on it yet.
- ibara’s own speed per call:
computer_actmedian 2.5 s, 95th percentile 8.9 s, over 927 calls;browser_actmedian 1.4 s, 95th percentile 7.2 s, over 388 calls. - Approvals: 54 were asked; 52 were approved by the test runner, none denied, and 2 were questions to the person.
- The run stopped at 06:32 with 99 queued runs not started, so not every task got two runs per setup on each machine. Sample sizes are uneven.
| Agent and model (provider) | Passed |
|---|---|
| Claude Code 2.1.284, claude-opus-5-5 (Anthropic) | 22 of 23 |
| Codex 0.158.0, gpt-6-astra, reasoning low (OpenAI) | 22 of 23 |
| Codex 0.158.0, gpt-6-sol, reasoning low (OpenAI) | 21 of 23 |
| Claude Code 2.1.284, claude-sonnet-5 (Anthropic) | 18 of 23 |
| omp 18.4.3, grok-4.7 (xAI) | 18 of 22 |
| Codex 0.158.0, deepseek-v4.1-flash through OpenRouter (DeepSeek) | 17 of 21 |
| opencode 1.18.31, big-pickle, free (OpenCode Zen) | 14 of 22 |
| Task | Passed |
|---|---|
| Write and save a note in the text editor | 13 of 13 |
| Rename and move files in the file manager | 3 of 13 |
| Run a command in a terminal and report its output | 12 of 13 |
| Find a setting and change it (editor line numbers) | 8 of 12 |
| Copy text from the terminal into the text editor | 11 of 11 |
| Find a file on the computer and deliver it here | 9 of 12 |
| Look up a fact on Wikipedia | 13 of 14 |
| Fill in and submit a sign-up form | 12 of 14 |
| Navigate a multi-page site and read a table | 14 of 14 |
| Search, download a file and deliver it here | 13 of 14 |
| Get past a cookie banner and a popup | 14 of 14 |
| Type accented and non-Latin text into a page | 10 of 13 |
0.1.0-18, focused rerun
On 29 September, from 06:42 to 07:43, the published 0.1.0-18 ran 5 of the tasks again on tater0 and tater1, with the same isolation and checks: the file manager, the editor setting, the Wikipedia fact, the sign-up form and the accented text. Three setups ran each task once on each machine: 30 runs, and every run counts. This is a smaller run than the one above, not a full benchmark.
- Claude Code with claude-opus-5-5 and Codex with gpt-6-astra passed all 20 of their runs. On 0.1.0-17 the same tasks and setups passed 17 of 18.
- The file manager task passed 4 of 4 for those two setups, on both machines. On 0.1.0-17 it passed 3 of 13 across all setups.
- All runs: 21 of 30 passed. 8 of the 9 failures were opencode with the free big-pickle model making no call at all, because the free model’s quota ran out; the 9th was that model’s malformed calls. None of the failures was ibara’s.
Tasks, first run#
On 28 September, real agents ran 6 desktop and 6 browser tasks on tater0 and tater1, each from a clean start: 73 runs in all, and every run counts. A run passes when every check the task declares passes at its finish, within its time limit. Time is the agent’s working time, including the model’s thinking, minus any wait for a person’s approval. Neither computer had a mouse plugged in. The runs used release candidate 0.1.0-9, and on tater1 from 12:30 a development build with the no-mouse fix. Of the 24 failed runs, 17 failed for ibara defects that were found in this run and are fixed in the release.
| Task | Passed | Median time (range) | Median tool calls (range) |
|---|---|---|---|
| Write and save a note in the text editor | 9 of 9 | 71.7 s (53–135) | 13 (10–25) |
| Rename and move files in the file manager | 0 of 5 | 159.5 s (38–420) | 11 (9–65) |
| Run a command in a terminal and report its output | 5 of 5 | 44.5 s (37–48) | 9 (8–13) |
| Find a setting and change it (editor line numbers) | 2 of 3 | 110.0 s (92–162) | 25 (15–30) |
| Copy text from the terminal into the text editor | 2 of 3 | 89.0 s (78–93) | 15 (15–16) |
| Find a file on the computer and deliver it here | 1 of 3 | 63.5 s (50–108) | 17 (13–37) |
| Look up a fact on Wikipedia | 5 of 5 | 73.0 s (48–80) | 13 (10–17) |
| Fill in and submit a sign-up form | 8 of 14 | 129.1 s (78–300) | 30 (17–51) |
| Navigate a multi-page site and read a table | 5 of 5 | 102.0 s (71–122) | 19 (13–25) |
| Search, download a file and deliver it here | 2 of 5 | 195.1 s (148–252) | 41 (27–59) |
| Get past a cookie banner and a popup | 5 of 5 | 60.0 s (40–92) | 13 (8–20) |
| Type accented and non-Latin text into a page | 5 of 11 | 140.7 s (93–300) | 25 (16–62) |
The file manager couldn’t be opened on these builds; the release opens it.
| Agent and model (provider) | Passed | Median time (range) | Median tool calls (range) |
|---|---|---|---|
| Claude Code, claude-opus-5-5 (Anthropic) | 17 of 23 | 94.5 s (37–420) | 16 (8–62) |
| Claude Code, claude-sonnet-5 (Anthropic) | 3 of 4 | 152.0 s (81–209) | 35.5 (12–50) |
| Codex, gpt-6-astra, reasoning low (OpenAI) | 14 of 26 | 84.5 s (38–170) | 15 (8–31) |
| Codex, gpt-6-sol, reasoning low (OpenAI) | 4 of 4 | 114.5 s (89–171) | 18.5 (13–33) |
| Codex, deepseek-v4.1-flash through OpenRouter (DeepSeek) | 2 of 2 | 82.2 s (72–93) | 22 (19–25) |
| omp, grok-4.7 (xAI) | 7 of 10 | 164.1 s (44–420) | 30 (13–65) |
| opencode, big-pickle, free (OpenCode Zen) | 2 of 4 | 230.4 s (122–300) | 43 (25–51) |
Versions: Claude Code 2.1.283, Codex 0.156.1, omp 18.3.5, opencode 1.18.31. Google Gemini wasn’t tested: its only route had reached its usage limit. Sample sizes are uneven.
Other figures from the same run:
- ibara’s own speed per call, as the agent’s harness saw it, including the network hop and not counting model time:
computer_act(up to 8 steps) median 2.0 s, 95th percentile 6.4 s, over 367 calls;browser_actmedian 0.67 s, 95th percentile 7.4 s, over 140 calls. - Clean finish: 40 of 73 runs said the task was complete and left no window open. 68 of 73 left no window open.
- Approvals: 39 were asked in 21 runs. 35 were approved, none denied and 4 unanswered. The median from asking to answering was 4.0 s; almost all were answered by the test runner, not a person.
- No mouse, before and after the fix (tater1 browser tasks): 5 of 8 passed with 12 refused pointer inputs on 0.1.0-9, and 17 of 25 passed with none on the fixed build.
Earlier acceptance runs
| Task | Result | Machines | Date |
|---|---|---|---|
| Fill in and submit a web form in a signed-in browser: typing, choosing an option, ticking a box, and clicking a button that only responds to real clicks | Passed in 14.2–14.6 s, 10 tool calls (a scripted client, no model) | tater0, tater1 | |
| Make a brand-new browser profile ready for agent control, with no clicks | Ready in 2.0–3.1 s | tater0, tater1 | |
| Open a text editor, type a line, save it and hand the file back | Passed in 11.4–11.5 s, 7 tool calls; 12.2 s on a machine limited to 2 GB | tater0, tater1 | |
| The same save-a-file task, 100 times in a row | 99 passed (ibara 0.1.0-9), median 12.6–12.9 s with no drift; the one failure was a monitor reconnecting while the text was being typed | tater0 | |
| Codex, Claude Code, opencode and omp each connect themselves from one prompt, then save a file on another computer | All 4 connected (ibara 0.1.0-8); 5 of 5 tasks completed across 3 model providers | tater1 |
Input accuracy#
On tater0, 28 September, ibara 0.1.0-9, with scripted clients (no model).
| Measure | Result | Machines | Date |
|---|---|---|---|
| Clicks land where aimed: page buttons in Chrome at scale 1 and 1.5 and zoom 100 and 125 %, and a GTK 4 app at scale 1 and 1.5 | 300 of 300 hit first try, 0 px from the target’s center. An idle mouse was plugged in during this test. | tater0 | |
| Page button presses in the click test above | Every one of 200 arrived as real input | tater0 | |
| Your own mouse while an agent works | Your pointer moves 17 ms after you move the mouse (median, 15 of 15), and the agent’s next click went through each time | tater0 |
Handoff#
| Measure | Result | Machines | Date |
|---|---|---|---|
| Take Control, from choosing it in the console to the viewer on screen (20 tries on Wi-Fi; tater0 on 0.1.0-9, tater1’s console on a development build of it) | 8.2 s median (8.1–10.5 s) | tater1 viewing tater0 | |
| Your first key, after the viewer shows | Arrives 50 ms later | tater1 viewing tater0 | |
| Hand Back, until the computer accepts an agent again (40 tries) | 0.7 s median (0.6–0.9 s). The viewer is gone in 0.54 s and the stream stops in 0.62 s. | tater1 viewing tater0 | |
| A viewer whose access is revoked is disconnected | In 0.59 s | tater0 |
Footprint#
Memory and CPU of every ibara process, plus the console’s share of the Omarchy shell, on tater0 with its full 7.6 GB of memory, 28 September, ibara 0.1.0-9. The targets are ibara’s own limits.
| State | Memory | CPU (one core) | Target |
|---|---|---|---|
| Idle, console not opened | 17.8 MiB | 0.4 % | 40 MiB |
| Idle, console open on the fleet | 58.2 MiB | 0.9 % | 40 MiB (over) |
| Agent working, console closed | 117.0 MiB (96.9–157.0) | 57 % | 120 MiB |
Someone has control of this computer (ibara-stream) | 71.3 MiB | 38 % | 100 MiB |
You have control of another computer (ibara-view, measured on tater1) | 80 MiB average, 105 MiB peak | 20 % | 150 MiB |
On a machine limited to 2 GB of RAM (ibara 0.1.0-7 on tater1, one run per state):
| State | Memory | Target | Date |
|---|---|---|---|
| Idle, console not opened | 18.6 MiB | 40 MiB | |
| Agent working, console closed | 121.6 MiB | 120 MiB | |
| Idle after every part of the console has been opened | 48.8 MiB | 40 MiB | |
| Agent working with the console open | 187.1 MiB | 120 MiB |
- Network, sent by tater0: watching a still desktop in the Screen tab used 24.5 kbit/s, against 1.0 kbit/s without watching (ibara 0.1.0-10). Take Control used about 1,021 kbit/s on a still screen and on a moving one (0.1.0-9). Watch on a busy screen hasn’t been measured on the test machines since 0.1.0-10.
- Disk, on tater0 (ibara 0.1.0-9): 137.9 MiB of packages (ibara 10.5, ibara-stream 23.0, ibara-view 2.2 and Cua’s driver 102.3). In 0.1.0-17, ibara’s own 3 packages are a 15.5 MB download and 36.4 MiB installed. ibara’s data grows with use: 592 MiB after 790 tasks on tater0.
- Previews: on a test wall of 100 made-up computers, ibara kept its stored previews at or below 63.9 MiB, within its 64 MiB cap (development build, 24 September).
Power hasn’t been measured yet. The figures on tested hardware are ratings and typical draw.
Method#
The 29 September run on the published 0.1.0-17 is the latest full benchmark, and the 0.1.0-18 rerun repeats 5 of its tasks. The 28 September tables are the first full benchmark, on earlier builds; the acceptance runs beneath them are earlier and smaller. Every run counts, failures included, and nothing is dropped. Each result states the machine, the Omarchy version, the ibara build, and the agent and model that ran it.
| Benchmark | What it measures | How it’s run |
|---|---|---|
| Desktop tasks | Completion rate and time for everyday tasks across apps | A fixed set of tasks, each from a clean start. A run passes when every check the task declares passes at its finish. |
| Browser tasks | Completion rate and time for common browser flows | Forms, navigation and search on signed-in browsers, scored the same way. |
| Input accuracy | How often a click or keystroke lands where intended the first time | Targets at known positions, across screen scales and zoom levels; distance off target for every click. |
| Handoff | Time from Take Control to your input, and from Hand Back to available | Timed from the choice in the console to the viewer on screen and the first input that arrives, and to the computer accepting an agent again. |
| Footprint | Memory, CPU and network, idle and while an agent works | Each state on each tested machine. Power hasn’t been measured yet. |