BENCHMARKS

Benchmarks.

Measured results for computer use and browser use on real, owned hardware: how long tasks take, where clicks land, how fast control changes hands, and how much memory ibara uses. Each number has its date and machines, and the method says how the full benchmarks are run.

Updated

Measured so far#

These numbers come from runs on the tested hardware, tater0 and tater1: five-year-old mini PCs with 10-watt Celerons. Both ran Omarchy 4.0.4 with Hyprland 0.56.2; Cua’s driver was 0.29.1 on tater0 and 0.28.2 on tater1. On 28 September real agents ran the first full benchmark, and the handoff, input and memory measurements were repeated. On 29 September the same tasks ran again on the published release, 0.1.0-17. The earlier acceptance runs stay below, dated, and the acceptance runs are also listed, passed or not, in the test results.

Tasks on the release#

On 29 September, from 02:56 to 06:32, real agents ran the same 12 tasks and checks as on 28 September on the published package, ibara 0.1.0-17, the same on both machines. tater0 had Google Chrome and a wireless mouse receiver plugged in; tater1 had Chromium, no mouse, and was freshly installed from GitHub that night. Each agent had only the ibara MCP server, the packaged ibara skill and the ibara note in its instructions; the project’s own source and notes were hidden from it. Every run counts.

  • 132 of 157 runs passed (84 %), against 49 of 73 (67 %) on 28 September. Desktop tasks 56 of 74, browser tasks 76 of 83; 18 runs hit the time limit. tater0 67 of 81, tater1 65 of 76.
  • Failures: 13 were ibara’s, 8 were the agents’, and 4 were the benchmark’s own setup, fixed during the run and still counted. The weakest task, the file manager (3 of 13), failed because the rename popup took the keyboard and refused every input; that caused 10 of the 13 ibara failures. Some clicks on its buttons were also refused along the way, which cost the agents extra calls. The other 3 ibara failures came from the editor keeping its settings where the agent’s usual tools can’t see them; since 0.1.0-20 the editor uses your own Mousepad settings. ibara also sometimes marked finished browser tasks incomplete, because its text check read only part of the page. 0.1.0-18, released the same morning, fixes the popup, the refused clicks and the text check; a focused rerun of 5 tasks on it is below, and the full benchmark hasn’t been run on it yet.
  • ibara’s own speed per call: computer_act median 2.5 s, 95th percentile 8.9 s, over 927 calls; browser_act median 1.4 s, 95th percentile 7.2 s, over 388 calls.
  • Approvals: 54 were asked; 52 were approved by the test runner, none denied, and 2 were questions to the person.
  • The run stopped at 06:32 with 99 queued runs not started, so not every task got two runs per setup on each machine. Sample sizes are uneven.
Agent and model (provider)Passed
Claude Code 2.1.284, claude-opus-5-5 (Anthropic)22 of 23
Codex 0.158.0, gpt-6-astra, reasoning low (OpenAI)22 of 23
Codex 0.158.0, gpt-6-sol, reasoning low (OpenAI)21 of 23
Claude Code 2.1.284, claude-sonnet-5 (Anthropic)18 of 23
omp 18.4.3, grok-4.7 (xAI)18 of 22
Codex 0.158.0, deepseek-v4.1-flash through OpenRouter (DeepSeek)17 of 21
opencode 1.18.31, big-pickle, free (OpenCode Zen)14 of 22
TaskPassed
Write and save a note in the text editor13 of 13
Rename and move files in the file manager3 of 13
Run a command in a terminal and report its output12 of 13
Find a setting and change it (editor line numbers)8 of 12
Copy text from the terminal into the text editor11 of 11
Find a file on the computer and deliver it here9 of 12
Look up a fact on Wikipedia13 of 14
Fill in and submit a sign-up form12 of 14
Navigate a multi-page site and read a table14 of 14
Search, download a file and deliver it here13 of 14
Get past a cookie banner and a popup14 of 14
Type accented and non-Latin text into a page10 of 13

0.1.0-18, focused rerun

On 29 September, from 06:42 to 07:43, the published 0.1.0-18 ran 5 of the tasks again on tater0 and tater1, with the same isolation and checks: the file manager, the editor setting, the Wikipedia fact, the sign-up form and the accented text. Three setups ran each task once on each machine: 30 runs, and every run counts. This is a smaller run than the one above, not a full benchmark.

  • Claude Code with claude-opus-5-5 and Codex with gpt-6-astra passed all 20 of their runs. On 0.1.0-17 the same tasks and setups passed 17 of 18.
  • The file manager task passed 4 of 4 for those two setups, on both machines. On 0.1.0-17 it passed 3 of 13 across all setups.
  • All runs: 21 of 30 passed. 8 of the 9 failures were opencode with the free big-pickle model making no call at all, because the free model’s quota ran out; the 9th was that model’s malformed calls. None of the failures was ibara’s.

Tasks, first run#

On 28 September, real agents ran 6 desktop and 6 browser tasks on tater0 and tater1, each from a clean start: 73 runs in all, and every run counts. A run passes when every check the task declares passes at its finish, within its time limit. Time is the agent’s working time, including the model’s thinking, minus any wait for a person’s approval. Neither computer had a mouse plugged in. The runs used release candidate 0.1.0-9, and on tater1 from 12:30 a development build with the no-mouse fix. Of the 24 failed runs, 17 failed for ibara defects that were found in this run and are fixed in the release.

TaskPassedMedian time (range)Median tool calls (range)
Write and save a note in the text editor9 of 971.7 s (53–135)13 (10–25)
Rename and move files in the file manager0 of 5159.5 s (38–420)11 (9–65)
Run a command in a terminal and report its output5 of 544.5 s (37–48)9 (8–13)
Find a setting and change it (editor line numbers)2 of 3110.0 s (92–162)25 (15–30)
Copy text from the terminal into the text editor2 of 389.0 s (78–93)15 (15–16)
Find a file on the computer and deliver it here1 of 363.5 s (50–108)17 (13–37)
Look up a fact on Wikipedia5 of 573.0 s (48–80)13 (10–17)
Fill in and submit a sign-up form8 of 14129.1 s (78–300)30 (17–51)
Navigate a multi-page site and read a table5 of 5102.0 s (71–122)19 (13–25)
Search, download a file and deliver it here2 of 5195.1 s (148–252)41 (27–59)
Get past a cookie banner and a popup5 of 560.0 s (40–92)13 (8–20)
Type accented and non-Latin text into a page5 of 11140.7 s (93–300)25 (16–62)

The file manager couldn’t be opened on these builds; the release opens it.

Agent and model (provider)PassedMedian time (range)Median tool calls (range)
Claude Code, claude-opus-5-5 (Anthropic)17 of 2394.5 s (37–420)16 (8–62)
Claude Code, claude-sonnet-5 (Anthropic)3 of 4152.0 s (81–209)35.5 (12–50)
Codex, gpt-6-astra, reasoning low (OpenAI)14 of 2684.5 s (38–170)15 (8–31)
Codex, gpt-6-sol, reasoning low (OpenAI)4 of 4114.5 s (89–171)18.5 (13–33)
Codex, deepseek-v4.1-flash through OpenRouter (DeepSeek)2 of 282.2 s (72–93)22 (19–25)
omp, grok-4.7 (xAI)7 of 10164.1 s (44–420)30 (13–65)
opencode, big-pickle, free (OpenCode Zen)2 of 4230.4 s (122–300)43 (25–51)

Versions: Claude Code 2.1.283, Codex 0.156.1, omp 18.3.5, opencode 1.18.31. Google Gemini wasn’t tested: its only route had reached its usage limit. Sample sizes are uneven.

Other figures from the same run:

  • ibara’s own speed per call, as the agent’s harness saw it, including the network hop and not counting model time: computer_act (up to 8 steps) median 2.0 s, 95th percentile 6.4 s, over 367 calls; browser_act median 0.67 s, 95th percentile 7.4 s, over 140 calls.
  • Clean finish: 40 of 73 runs said the task was complete and left no window open. 68 of 73 left no window open.
  • Approvals: 39 were asked in 21 runs. 35 were approved, none denied and 4 unanswered. The median from asking to answering was 4.0 s; almost all were answered by the test runner, not a person.
  • No mouse, before and after the fix (tater1 browser tasks): 5 of 8 passed with 12 refused pointer inputs on 0.1.0-9, and 17 of 25 passed with none on the fixed build.

Earlier acceptance runs

TaskResultMachinesDate
Fill in and submit a web form in a signed-in browser: typing, choosing an option, ticking a box, and clicking a button that only responds to real clicksPassed in 14.2–14.6 s, 10 tool calls (a scripted client, no model)tater0, tater1
Make a brand-new browser profile ready for agent control, with no clicksReady in 2.0–3.1 stater0, tater1
Open a text editor, type a line, save it and hand the file backPassed in 11.4–11.5 s, 7 tool calls; 12.2 s on a machine limited to 2 GBtater0, tater1
The same save-a-file task, 100 times in a row99 passed (ibara 0.1.0-9), median 12.6–12.9 s with no drift; the one failure was a monitor reconnecting while the text was being typedtater0
Codex, Claude Code, opencode and omp each connect themselves from one prompt, then save a file on another computerAll 4 connected (ibara 0.1.0-8); 5 of 5 tasks completed across 3 model providerstater1

Input accuracy#

On tater0, 28 September, ibara 0.1.0-9, with scripted clients (no model).

MeasureResultMachinesDate
Clicks land where aimed: page buttons in Chrome at scale 1 and 1.5 and zoom 100 and 125 %, and a GTK 4 app at scale 1 and 1.5300 of 300 hit first try, 0 px from the target’s center. An idle mouse was plugged in during this test.tater0
Page button presses in the click test aboveEvery one of 200 arrived as real inputtater0
Your own mouse while an agent worksYour pointer moves 17 ms after you move the mouse (median, 15 of 15), and the agent’s next click went through each timetater0

Handoff#

MeasureResultMachinesDate
Take Control, from choosing it in the console to the viewer on screen (20 tries on Wi-Fi; tater0 on 0.1.0-9, tater1’s console on a development build of it)8.2 s median (8.1–10.5 s)tater1 viewing tater0
Your first key, after the viewer showsArrives 50 ms latertater1 viewing tater0
Hand Back, until the computer accepts an agent again (40 tries)0.7 s median (0.6–0.9 s). The viewer is gone in 0.54 s and the stream stops in 0.62 s.tater1 viewing tater0
A viewer whose access is revoked is disconnectedIn 0.59 stater0

Footprint#

Memory and CPU of every ibara process, plus the console’s share of the Omarchy shell, on tater0 with its full 7.6 GB of memory, 28 September, ibara 0.1.0-9. The targets are ibara’s own limits.

StateMemoryCPU (one core)Target
Idle, console not opened17.8 MiB0.4 %40 MiB
Idle, console open on the fleet58.2 MiB0.9 %40 MiB (over)
Agent working, console closed117.0 MiB (96.9–157.0)57 %120 MiB
Someone has control of this computer (ibara-stream)71.3 MiB38 %100 MiB
You have control of another computer (ibara-view, measured on tater1)80 MiB average, 105 MiB peak20 %150 MiB

On a machine limited to 2 GB of RAM (ibara 0.1.0-7 on tater1, one run per state):

StateMemoryTargetDate
Idle, console not opened18.6 MiB40 MiB
Agent working, console closed121.6 MiB120 MiB
Idle after every part of the console has been opened48.8 MiB40 MiB
Agent working with the console open187.1 MiB120 MiB
  • Network, sent by tater0: watching a still desktop in the Screen tab used 24.5 kbit/s, against 1.0 kbit/s without watching (ibara 0.1.0-10). Take Control used about 1,021 kbit/s on a still screen and on a moving one (0.1.0-9). Watch on a busy screen hasn’t been measured on the test machines since 0.1.0-10.
  • Disk, on tater0 (ibara 0.1.0-9): 137.9 MiB of packages (ibara 10.5, ibara-stream 23.0, ibara-view 2.2 and Cua’s driver 102.3). In 0.1.0-17, ibara’s own 3 packages are a 15.5 MB download and 36.4 MiB installed. ibara’s data grows with use: 592 MiB after 790 tasks on tater0.
  • Previews: on a test wall of 100 made-up computers, ibara kept its stored previews at or below 63.9 MiB, within its 64 MiB cap (development build, 24 September).

Power hasn’t been measured yet. The figures on tested hardware are ratings and typical draw.

Method#

The 29 September run on the published 0.1.0-17 is the latest full benchmark, and the 0.1.0-18 rerun repeats 5 of its tasks. The 28 September tables are the first full benchmark, on earlier builds; the acceptance runs beneath them are earlier and smaller. Every run counts, failures included, and nothing is dropped. Each result states the machine, the Omarchy version, the ibara build, and the agent and model that ran it.

BenchmarkWhat it measuresHow it’s run
Desktop tasksCompletion rate and time for everyday tasks across appsA fixed set of tasks, each from a clean start. A run passes when every check the task declares passes at its finish.
Browser tasksCompletion rate and time for common browser flowsForms, navigation and search on signed-in browsers, scored the same way.
Input accuracyHow often a click or keystroke lands where intended the first timeTargets at known positions, across screen scales and zoom levels; distance off target for every click.
HandoffTime from Take Control to your input, and from Hand Back to availableTimed from the choice in the console to the viewer on screen and the first input that arrives, and to the computer accepting an agent again.
FootprintMemory, CPU and network, idle and while an agent worksEach state on each tested machine. Power hasn’t been measured yet.