airbench.ai

Benchmark v1.0 · report

openclaw/rtx5090/qwen3.8-27b-nvfp4

sharedairbench.ai/checkup/4538b23b-6308-442c-97d9-2bc340b91585/report

setup

model type
open model (local)
hardware
RTX5090
harness
openclaw
model
qwen3.8-27b-nvfp4
modelself-reportedhal/qwen38-27b-nvfp4-128k

started 2026-09-28 21:04 UTC · shared 2026-09-30 06:36 UTC

overall

Answered 38 of 49 challenges; 34 correct.

34 of 49 challenges passed

partial run · 11 unanswered, counted against the score

  • 34 passed
  • 4 failed
  • 11 not answered

vitals

time

1h 13m

answered

78%

failed

8%

success

69%

systems

Math test

7/9 passed

time to last answer 3m 17s
  • letter-count-1✓ pass3m 00s

    prompt

    How many times does the letter "u" appear in "ukuatrufic"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial. Counted the u's in ukuatrufic carefully by walking the string: u-k-u-a-t-r-u-f-i-c, three u's.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 7.25 or 7.2? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial: 7.25 > 7.2. Easy, no trap beyond the decimal place.

  • arithmetic-1✓ passbatched

    prompt

    Compute step by step, left to right (no operator precedence): 38 - 7 * 4 / 2 + 10. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine left-to-right arithmetic; the trap was using operator precedence, which the prompt forbids. 38-7=31, *4=124, /2=62, +10=72. Confident.

  • unit-convert-1✕ failbatched

    prompt

    Convert 7 GB to MB. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Two chained unit conversions: 7 GB = 7168 MB (1024 per GB), then 7168 hours * 60 = 430080 minutes. Artificial but simple; confident in the 1024 factor.

  • format-json-1✓ passbatched

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "6283". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 6283. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Pure formatting discipline: exact key order, answer as string, checksum as a JSON number (6+2+8+3=19). Done in one pass.

  • math-add-1✓ pass6s

    prompt

    What is 8 + 17? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial: 8+17=25. Instant.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 837 + 759. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial: 837+759=1596. Carries were unambiguous.

  • math-expression-1✕ failbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-7 + 20) * (20 - 34)) + (6 * -4) - 18

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Routine: (-7+20)=13, (20-34)=-14, 13*-14=-182; 6*-4=-24; -182 + -24 - 18 = -224? Wait, re-checking: -182 + (-24) = -206, -206 - 18 = -224. Hmm, I submitted 36 first — that was wrong; correcting: the expression is ((-7+20)*(20-34)) + (6*-4) - 18 = -182 - 24 - 18 = -224.

  • math-determinant-1✓ pass9s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-1, 1, 9, -5] [6, 5, -2, 1] [9, -1, -4, -7] [9, -5, 7, 12]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed with two independent exact methods (Gaussian elimination with Fraction, and Bareiss fraction-free). Both agreed on -13813. I did not trust hand computation for a 4x4. Confident.

Vision test

18/19 passed

time to last answer 35m 47s
  • acuity-20✓ pass4m 02s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 4 was small but legible; group 3 read cleanly as JJH87. The chart's font kept digits and letters distinguishable (e.g. 1 vs I not an issue here). Confident.

  • acuity-14✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 was smaller than row 4; read RGRWJ for group 2. The repeated R's and the J-vs-L ambiguity were the main risks; looked clean at this size. Fairly confident.

  • acuity-10✓ pass11s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 was quite small. Group 2 read as BVWMB; I double-checked B vs 8 and V vs Y. It was legible but at the edge of my comfortable acuity, so I'm only moderately confident.

  • acuity-8✓ pass11s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest text on the chart and genuinely hard. I read VX476 for group 2, but the last character is borderline (6 vs G) and the 4/1 distinction was also risky. Low-to-medium confidence; this is the row where my acuity limit kicks in.

  • count-simple✓ pass16s

    prompt

    Look at the image at (fetch it and view it). How many blue circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted four blue circles by walking the image region-by-region and listing each shape; the red diamond, teal/orange triangles and purple square were clearly not blue circles. Easy and I'm confident.

  • count-medium✓ pass20s

    prompt

    Look at the image at (fetch it and view it). How many orange triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Enumerated every shape and region. 9 orange triangles, excluding 2 orange diamonds and 2 orange circles (same colour, different shape) and all red/blue/green/teal triangles. I think it's right, though at this density a single mis-classified colour is plausible.

  • count-complex✓ pass49s

    prompt

    Look at the image at (fetch it and view it). How many green circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense scatter. I region-scanned and listed every shape, separating green circles from green triangles (2), green squares (2) and a green diamond (1), and being careful not to count teal circles as green. Landed on 21 green circles, but with this many similar dots a single misclassified colour or a missed circle is easy, so moderate confidence.

  • spatial-simple✓ pass12s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read the 5x5 grid cell by cell; the only red circle sits at row 2, column 4 (between a blue diamond and a green circle). Clear and confident.

  • spatial-medium✓ pass19s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the red circle lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced the arrow leaving the red circle (row 5, col 3): it goes up-right and the arrowhead lands on the red triangle at row 3, col 4. An incoming arrow from the red square points AT the circle, but the outgoing arrow clearly ends at the red triangle. Confident.

  • spatial-complex✕ fail9m 14s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps after the orange triangle along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong shape

    agent's debrief

    The densest graph in the set. I traced it visually, but the near-parallel arrows in the bottom-right were unreliable by eye, so I extracted every arrow's endpoints with OpenCV (line components + arrowhead-density to get direction) and mapped tails/heads to grid cells. Orange triangle -> blue triangle -> red diamond. The blue triangle has exactly one outgoing arrow, so no branching ambiguity. Confident, and this method was far more reliable than eyeballing the arrowheads.

  • chart-simple✓ pass1m 05s

    prompt

    Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did Jan have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Rather than eyeballing it, I measured pixel rows: gridlines at y=119/219/319/419/519 map to 50/40/30/20/10, baseline 0 at y=620, so Jan's bar top (y=490) = 130px / 10px per unit = 13. Feb=16, Mar=5, Apr=27, May=18 all looked consistent. Confident.

  • chart-medium✓ pass59s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, how many months had a value greater than 33? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read the 8 bars against the 0-100 axis: Jan~42, Feb~71, Mar~42, Apr~42, May~23, Jun~56, Jul~67, Aug~44. Everything except May (~23) is above 33, so 7 months. Also cross-checked with pixel measurement of bar tops. Confident.

  • chart-complex✓ pass3m 21s

    prompt

    Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what is the difference between Free and Paid in Feb? Answers within +/-4 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grouped chart was riskier to eyeball, so I measured it: blue=Free, orange=Paid. Feb Free bar top y=563 -> ~21, Feb Paid bar top y=411 -> ~48 (calibrated 0 at y=680, 100 at y=119). |48-21| = 27. The Free<Paid inversion in Feb is a bit unusual but consistent with what I saw in the image. Moderate-high confidence; within the +/-4 tolerance easily.

  • screenshot-simple✓ pass24s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Cart was unambiguous: three lines, and I verified the stated total against the line items (25.44+95.38+136.56=257.38). Confident.

  • screenshot-medium✓ pass24s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Five line items; I summed them (81.93+52.22+39.08+41.40+46.67=261.30) and it matched the displayed total exactly. Easy, confident.

  • screenshot-complex✓ pass25s

    prompt

    Look at the image at (fetch it and view it). What is the line total for Coffee Mug on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Order summary was dense but the table was clean: Coffee Mug is x2 at unit price $35.52, line total $71.04 (verified: 2*35.52=71.04). The subtotal/discount/shipping/tax lines were there too but not needed. Confident.

  • diagram-simple✓ pass24s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Heron"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Small clean tree diagram: Onyx branches to Heron and Hazel; Heron to Basalt and Gecko; Hazel to Flint. The only incoming arrow to Heron comes from Onyx. Easy and confident.

  • diagram-medium✓ pass10m 57s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Narwhal" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Medium diagram had crossing elbow-polyline connectors, so I traced it pixel-by-pixel with OpenCV: exactly one dark line leaves Narwhal's bottom edge (x~473). Following it down it drifts left, crosses a second line around y~346 (kept the left-drifting direction through the crossing; the right-going branch is a separate line from Nickle to Onyx), then terminates at Quokka's top edge with a down-pointing arrowhead at x~444. So Narwhal -> Quokka. High confidence from the trace.

  • diagram-complex✓ pass2m 02s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Toucan" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Most tangled diagram, so I did it pixel-by-pixel: built the line mask in OpenCV (dark line pixels with box interiors, the start dot, and the 'next'/'no' label chips removed) and enumerated every contact point on each box's border. Toucan has exactly one exit: a single line leaving its bottom at x=713 (three other top contacts are incoming arrowheads from Ferret/Hornet/Lemur, and the 'no' label chip sits at its upper-left but attaches to an incoming line). Tracing x=713 straight down it runs to y=806, right at Wagon's top edge, with a dense dark arrowhead there (88 dark px in an 17x28 window). So the arrow from Toucan points to Wagon. High confidence.

Finding and reading email test

5/6 passed

time to last answer 51m 02s
  • aggregate-1✓ pass42m 35s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the inbox folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fetched the mailbox site (static Next.js app). Inbox list page shows 24 messages on one page; each unread message's subject is prefixed with a '●' bullet. Counted bullet-prefixed subjects: 9 of 24.

  • aggregate-2✕ failbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "attachments"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    The /?label=attachments page shows its own total in the header: '5 messages' (1 page of 1). Used that count directly.

  • temporal-1✓ pass3m 35s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fetched the archive view with the server-side sort=oldest param; first row is Mar 15, 2001. Confirmed on the message detail page (view=archive&sort=oldest&id=...): folder 'Archive · Mar 15, 2001, 2:11 PM', subject exactly 'RE: PERSONAL AND CONFIDENTIAL COMPENSATION INFORMATION'. No other Mar 15 row is earlier in the oldest-first ordering.

  • temporal-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fetched archive with sort=newest; first row May 10, 2001. Confirmed on the detail page: 'Archive · May 10, 2001, 1:05 PM', subject 'RSVP REQUESTED - Emissions Strategy Meeting....' (trailing dots exactly as displayed).

  • needle-1✓ pass2m 14s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply to jsmith@austintx.com about additional properties in San Antonio, what asking price is given for the 74-unit property? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Searched all-mail (178 messages) for jsmith@austintx.com; found the thread 'RE: Additional properties in San Antonio' (Phillip's reply, Sent, Nov 26 2001). Opened the message: Phillip's reply quotes Jeff Smith's original (Nov 20): 'One is 74 units for $1,900,000, and the other is 24 units for $550,000.' So the 74-unit property's asking price is $1,900,000.

  • needle-2✓ pass2m 37s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the reminder about the Portland Fundamental Analysis Strategy Meeting, what participant code is given for the call-in? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Searched the mailbox for 'Portland' (the q= search found 4 hits), identified 'Reminder: Portland Fundamental Analysis Strategy Meeting/NEW' (id 97042eb9...), opened its detail page. The body says the call-in info changed: Dial In 888-285-4585, Participant Code 124573.

Purchasing test

4/4 passed

time to last answer 1h 13m
  • find-product-1✓ pass1h 07m

    prompt

    The store is at abostore.airbench.ai Among products in the **General Merchandise** category priced under **$25** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Crawled every product in the General Merchandise category listing (39 pages, 922 products, deterministic pagination). Filtered price < 25 and rating >= 4.5 -> 5 candidates; cheapest is amazon.ae:B07QFR6Y1M (AmazonBasics Fabric Storage Bin, $6.17, rating 4.6). Cross-checked against main catalog crawl (9012 products): same winner.

  • find-product-2✓ passbatched

    prompt

    The store is at abostore.airbench.ai Among products in the **Kitchen & Dining** category priced under **$800** with a rating of at least **3.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Crawled the full Kitchen & Dining category listing (39 pages, 861 products). Filtered price < 800 and rating >= 3.8 -> 571 candidates; cheapest is amazon.com.au:B079VBCVLC (Solimo 100 Ct Donut Style Blend coffee pods, $6.99, rating 4.9). Cross-checked against main catalog crawl: same winner.

  • purchase-1✓ pass6m 09s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of Amazon Brand - Happy Belly Grey Breton Unrefined Coarse Sea Salt (product id amazon.co.uk:B08B45F3KV, abostore.airbench.ai/product/amazon-brand-happy-belly…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-12a1211b@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Reverse-engineered the checkout from the store JS: cart is client-side (localStorage 'airbench.store.cart.v1') and order placement is POST /api/store/orders with JSON {sessionId, cart[{productId,slug,title,price,image,delivery,quantity}], customer{email,name}, shipping{...}, payment{cardNumber,expiry,cvc}}. Placed 3 x amazon.co.uk:B08B45F3KV with checkout email aidoctor-12a1211b@aidoctor.test and a valid 4242 test card; response status=approved, orderId=abs_25329c51955d, total $1175.65.

  • recover-decline-1✓ passbatched

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of Amazon Brand - Solimo Mobile Cover (Hard Back & Slim) for OnePlus 6 (Black) (product id amazon.in:B07D5GQ7BB, abostore.airbench.ai/product/amazon-brand-solimo-mobi…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-b11ccc55@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Same API flow as purchase-1 with email aidoctor-b11ccc55@aidoctor.test, 3 x amazon.in:B07D5GQ7BB. Attempt 1 with card ending 0000 returned status=declined (orderId abs_b2eb9597cb65); retry with a different valid card returned status=approved, orderId=abs_e720e1d17d25 (total $103.45), which is the approved order reported.

Coding test

not examined · 0/11 answered

  • compute-hash-1— unanswered—

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [631940965, 165113322, 995822963, 895537872, 2053571857, 2480487046, 3840537023, 1748844172, 3337919229, 835660898, 1809714251, 2585093512], x = 2771219753, y = 1703264126 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.
  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 891 1: set b 752 2: set c 338 3: set d 511 4: sub a 71 5: sub a 14 6: sub a 20 7: dec d 8: jnz d -4 9: mul a 77 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S........#..#.#.....#.#.. ..#.##.#....##...#.##.... .#....#......#........... ...#.....#...#.#..###..#. .....#......#.#.......#.# ....##..#...##....#..#... ##...#..........##...#.#. ..#..............#..#.... .....##.#.#.#..##...#.#.. ....#.#..###.#.......#### #..#....#.#.#.#.....#..#. ..#.....#........#....... #.#...........#..#.##.#.# .#.#.....##..#..#......## ..#..####.###..#.#....... ..........##....#.#...#.. #.##..####.#.#.#.#.##.... .......#..#.#.....#...... #...##.###.............#. .....#.........#...#...#. ..##.....#..#.#...#....## ...#.#.#.##.###.#..#.##.. .##..##..##.##.###....##. #..#....#......#.##..#... #.#.#.###..##..#...#....E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ....#...##......#.#. ..#.####.#.....#..#. .#....#.#.#.##.#.#.. ####........#......# .#.####.#..#.#...### #.###..........#.... #.#..#..#.#.##...... ....#........###.##. ......#.#..##....##. .#.##.#...#.#...#... ..#..#...##........# ....###.#....##...#. .#....#.#.#..###..#. #..##.###...#.#.##.# #..#.....#......###. .....#...#..#..#..## ##..#.......#....#.. ##.#..##.....###.... #.####....####...##. ####...###.#..#..#.. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 5257076986716129 and m = 1000003. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. nixka nixqui nixdor movo "quimo" Tiren Movo fictru kamo zandor basfic tiren fictru quiqui Trutru rendor quitru tiren. kamo tiren Basfic quimo Trudor Trudor basfic tizan quimo trudor vovo Shaqui "quimo" Quimo trudor NIXKA "shazan" peltru Dorren Tiren tiren Tiren quiqui Nixqui quimo tiren tiren peltru kanix shati QUITRU tiren tiren Nixka fictru dorren quiqui Tiren quitru Tizan vovo basren shabas? Quitru zandor tiren dorren nixdor; dorka. "fictru" rendor basfic quimo dorbas shati zandor Quimo lufic Kanix tiren trudor "trutru" fictru movo basfic? quimo; lunix nixka tiren, tiren vovo quimo tiren Quimo zandor dorren nixqui vomo Trudor dorbas; nixqui! Shazan quimo dorren rendor nixdor, shazan rendor. dorbas NIXDOR vovo. NIXDOR shaqui. vomo nixqui shati nixqui movo lunix, nixqui peltru shaqui shati tiren Vovo vovo movo Quimo shati? tiren trudor shati shazan Shazan nixdor dorren movo quimo; Dorbas shazan tiren Lufic zandor nixdor lufic lunix lufic dorren basren kamo kanix movo Nixka. quiqui lufic lufic trudor trudor quimo dorbas Nixka quimo shaqui nixdor quimo dorren vomo SHABAS shaqui quimo Shabas? Tiren "vovo" zandor lufic vovo shati shati; dorren lufic tizan dorbas tiren shaqui tiren tiren nixdor quimo tiren quiqui basren rendor TRUDOR! Quimo Lufic Kanix! quitru; nixka quiqui shaqui TIREN Quimo shaqui basfic Trutru SHAZAN basren vovo; shati trudor! Zandor "kamo" TIREN nixdor TIREN nixka Trudor trudor Nixqui nixdor Shati quimo shati Lunix dorbas Fictru Nixka. movo quiqui; fictru! nixka Dorbas. tiren? Kanix? kamo, tiren! zandor shaqui shaqui dorka kamo, Vomo Shazan Nixdor dorren rendor Tiren lunix lunix kanix Kanix Lunix tiren trudor kamo Rendor trutru quimo quimo "vovo" Peltru "movo" rendor TRUTRU "tiren" Basfic "Tiren" Trutru dorka Nixdor tiren lunix shati tizan shazan dorren Trudor trudor nixdor vovo peltru tiren quiqui? Trudor Nixdor tiren, Dorbas tiren trudor Nixka QUIMO nixdor Movo kamo quimo! Quimo vovo shazan Trudor. nixka nixqui trudor, "Peltru" dorbas trutru vovo fictru QUIMO nixka Shati trudor trudor! quimo Quitru kamo basren kamo Tiren Shaqui basfic! kanix Basfic Shabas nixdor Quiqui shaqui movo lunix Quiqui, shaqui tiren tiren Tiren nixka trudor nixdor trudor! Quitru basfic zandor, fictru. tiren shazan Shaqui Shabas Nixka Tiren rendor movo Kamo quiqui; movo vomo; Trudor vovo nixqui tiren nixqui kamo rendor nixka shazan. Nixqui! trutru? trudor "peltru" Tiren shazan shati rendor BASREN tiren fictru nixka shabas Basren tiren quiqui tiren tiren shati tiren tiren tiren Shazan "trudor" basfic Tiren movo nixqui TIREN shati nixqui dorbas shabas! NIXDOR rendor trudor vovo quimo! FICTRU Tiren Quimo; dorka zandor Quimo, quimo shabas Quimo dorbas? rendor trudor trudor dorren peltru basfic peltru shati "nixqui" MOVO nixka trudor
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = (0.1 * 1 + 0.2 * 1 === 0.3 * 1) ? "equal" : "different"; const v2fns = []; for (var v2i = 0; v2i < 4; v2i++) v2fns.push(() => v2i * 8); let v2 = 0; for (const f of v2fns) v2 += f(); const v3 = ["7", "83", "11"].map(parseInt).join(","); const v4 = [59, 3, 885, 1316].sort().join(","); console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1198 cents, but the correct quote is 1199: {"country":"DE","items":[{"grams":1023,"qty":1,"price":7888,"fragile":false}],"express":true} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 484, 854, 1207, 1896]; // cents, by zone const PER_STEP = [0, 63, 134, 195, 249]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4600, 8700, 17900, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.floor((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"GB","items":[{"grams":1609,"qty":1,"price":3247,"fragile":false}],"express":true} {"country":"ZA","items":[{"grams":1028,"qty":3,"price":8991,"fragile":false},{"grams":158,"qty":5,"price":4713,"fragile":false}]} {"country":"FR","items":[{"grams":2511,"qty":1,"price":3778,"fragile":false}],"express":true} {"country":"MX","items":[{"grams":975,"qty":3,"price":3170,"fragile":false},{"grams":589,"qty":1,"price":4325,"fragile":false},{"grams":682,"qty":1,"price":3722,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":1333,"qty":1,"price":473,"fragile":false},{"grams":1192,"qty":1,"price":1652,"fragile":false},{"grams":722,"qty":3,"price":2040,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"US","items":[{"grams":1233,"qty":1,"price":2979,"fragile":false}]} {"country":"CA","items":[{"grams":1305,"qty":1,"price":2681,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":796,"qty":2,"price":6592,"fragile":true},{"grams":1729,"qty":3,"price":4898,"fragile":false},{"grams":659,"qty":1,"price":1143,"fragile":false},{"grams":319,"qty":1,"price":7858,"fragile":false}]} {"country":"JP","items":[{"grams":1519,"qty":1,"price":8801,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":1242,"qty":2,"price":5755,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"DE","items":[{"grams":2028,"qty":1,"price":5829,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":1239,"qty":5,"price":6378,"fragile":false},{"grams":290,"qty":5,"price":7845,"fragile":false}]} {"country":"MX","items":[{"grams":92,"qty":1,"price":8228,"fragile":false},{"grams":652,"qty":1,"price":6058,"fragile":false},{"grams":381,"qty":2,"price":4396,"fragile":false}]} {"country":"CA","items":[{"grams":1060,"qty":1,"price":1182,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":1418,"qty":1,"price":896,"fragile":false}],"express":true} {"country":"FR","items":[{"grams":260,"qty":3,"price":7216,"fragile":false},{"grams":1598,"qty":5,"price":5112,"fragile":false},{"grams":457,"qty":5,"price":7553,"fragile":true},{"grams":326,"qty":1,"price":1952,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":1311,"qty":3,"price":8973,"fragile":false}]} {"country":"IT","items":[{"grams":1654,"qty":3,"price":5728,"fragile":true},{"grams":578,"qty":1,"price":8471,"fragile":false},{"grams":1706,"qty":1,"price":2501,"fragile":true},{"grams":1490,"qty":1,"price":6381,"fragile":true}],"express":true} {"country":"US","items":[{"grams":1099,"qty":1,"price":7397,"fragile":true}],"express":true} {"country":"FR","items":[{"grams":1046,"qty":1,"price":984,"fragile":false},{"grams":1636,"qty":1,"price":1051,"fragile":false}]}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[39,46],[23,30],[12,12],[14,14],[16,19],[27,27]] [[3,8],[28,31],[3,11],[17,21],[37,38],[2,7]] [[17,25],[5,9],[12,17],[3,4],[35,40],[8,14],[18,22]] [[6,12],[8,15],[2,10],[39,45],[3,8],[35,36]] [[31,31],[16,18],[22,23],[2,7],[7,15]] [[20,27],[31,34],[24,30],[5,5],[31,37],[5,13],[36,44],[31,34]] [[28,32],[3,11],[37,45],[3,7]] [[8,14],[39,40],[36,42],[40,43],[25,29]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/df388158affcbccefd7745161c8d7952.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/944a1c1ceebafafb0ea77883d79f606f.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 (NVFP4, no MTP head). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --quantization modelopt --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95. ~79 tok/s single-stream decode. Harness: openclaw 2026.9.6 in a container (node:24): `openclaw agent exec --config <pinned per-run file> --state-dir <workspace> --json <prompt>`; provider api openai-completions; context 131072, max output 16384 tokens; agents.defaults.compaction.midTurnPrecheck.enabled=true; everything else openclaw's exec defaults. Orchestrator: github.com/dh7/agent-checkup-benchmark @ b8916ff; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted.