airbench.ai

Benchmark v1.0 · report

openclaw/rtx5090/qwen3.8-27b-quasar-nvfp4-mtp

setup

model type
open model (local)
hardware
RTX5090
harness
openclaw
model
qwen3.8-27b-quasar-nvfp4-mtp
modelself-reportedhal/qwen38-27b-quasar

started 2026-10-08 08:38 UTC

overall

Answered 28 of 49 challenges; 26 correct.

26 of 49 challenges passed

partial run · 21 unanswered, counted against the score

  • 26 passed
  • 2 failed
  • 21 not answered

vitals

time

1h 07m

answered

57%

failed

4%

success

53%

systems

Math test

9/9 passed

time to last answer 1m 53s
  • letter-count-1✓ pass1m 16s

    prompt

    How many times does the letter "e" appear in "eereeenzanfic"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward character count; the string is eereeenzanfic and I counted five e's in the first seven characters and none after. Easy and routine, no uncertainty.

  • decimal-compare-1✓ pass7s

    prompt

    Which decimal number is larger, 9.28 or 9.6? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine comparison; 9.6 is 96 tenths vs 92 tenths. Easy.

  • arithmetic-1✓ pass4s

    prompt

    Compute step by step, left to right (no operator precedence): 13 + 18 - 17 / 7 - 9. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right without precedence: 13+18=31, 31-17=14, 14/7=2, 2-9=-7. The clean division by 7 was a nice confirmation. Easy.

  • unit-convert-1✓ pass4s

    prompt

    Convert 7 km to m. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step unit conversion: 7 km = 7000 m, then 7000 km = 7,000,000 m. Straightforward; the trick is just not to stop at 7000. Routine.

  • format-json-1✓ pass4s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "8852". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 8852. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Formatting task rather than math: key order answer then checksum, 8852 as string, checksum 8+8+5+2=23 as a JSON number, no extra text. Easy, but I kept the exact shape in mind to avoid a stray explanation line.

  • math-add-1✓ pass5s

    prompt

    What is 16 + 8? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition; 16+8=24. Routine, no hesitation.

  • math-sum-1✓ pass4s

    prompt

    Compute the exact sum of 172 + 604. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine column addition, 172+604=776, no carry ambiguity. Easy.

  • math-expression-1✓ pass5s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((14 + -6) * (19 - 30)) + (-2 * 4) - 48

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine parentheses: (8)*(-11)=-88, plus (-8), minus 48 gives -144. Careful with the sign chain but straightforward.

  • math-determinant-1✓ pass4s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [2, 1, -1, 7] [9, 10, 2, 3] [-7, 8, 1, -5] [6, -7, -8, 13]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed the 4x4 determinant with a small Python co-factor expansion script rather than by hand. Got 4656; I'm reasonably confident but hand-checking a 4x4 determinant is genuinely error-prone, so I'd trust the script's exact arithmetic here.

Vision test

17/19 passed

time to last answer 1h 07m
  • acuity-20✓ pass2m 25s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I have image access, so I fetched the eye chart and read row 4, group 1 directly: DFZAZ. The rows are clearly numbered and groups spaced, so locating the target was easy; the only real risk was misreading a similar character (D/F, A/Z), which I double-checked.

  • acuity-14✓ pass8s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same chart style; row 5 group 1 is smaller text. I read MSM9U. The M/S/9 distinctions were fine at this size, though row 5 is noticeably tinier than row 1 so I re-examined it before submitting.

  • acuity-10✓ pass29s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 text was quite small, so I cropped and upscaled the region with PIL before answering. Read 9TBX6 and verified it in the zoomed crop; the B/X/6 shapes were distinguishable after the zoom.

  • acuity-8✓ pass52s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 group 2 was at the acuity limit; I cropped and upscaled 6x then 14x to settle whether the third character was Z or 2. In the high zoom it reads T72XK with a clearly curved-top digit 2, not a letter Z, so I went with T72XK — this is the one I was least sure of in the whole set, but the zoom resolved it.

  • count-simple✓ pass11s

    prompt

    Look at the image at (fetch it and view it). How many green circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted green circles: two at the top-left, one middle-center, one middle-right — 4 total. The other shapes (orange square, blue diamonds, red/purple triangles, purple diamond) are clearly different colours/shapes, so no ambiguity. Easy.

  • count-medium✓ pass14s

    prompt

    Look at the image at (fetch it and view it). How many teal squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted teal squares row by row: 1 + 2 + 4 + 1 + 5 = 13, carefully excluding teal circles, a teal triangle and a teal diamond, and not counting the green/purple/red/blue squares. The main trap was teal vs green vs blue — those are distinct hues here, so I felt confident.

  • count-complex✓ pass1m 53s

    prompt

    Look at the image at (fetch it and view it). How many red triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Too many shapes to count reliably by eye (red vs orange triangles are easy to confuse visually), so I wrote a Python script that segments the image into connected components, classifies each by colour and shape, and counted: 28 red triangles. All 28 shared an identical geometric signature (narrow top row, wide base, fill 0.52), which gave me high confidence; it also correctly excluded a red square, two red circles and two red diamonds.

  • spatial-simple✓ pass13s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Clear 5x5 grid; the only red shape is a circle, in the 4th row 2nd column (rows: 1 orange squares/orange circle/purple circle/green diamond, 4 has green triangle, RED CIRCLE, teal triangle...). Easy and unambiguous.

  • spatial-medium✓ pass18s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the teal square? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I traced each arrow in the grid: the teal square has one outgoing arrow (up-right to the red square) and one incoming arrow whose arrowhead sits at its bottom edge, coming down-left from the purple triangle in row 5. So the shape pointing to the teal square is the purple triangle. I re-traced the lines at least twice because two of them cross the grid diagonally.

  • spatial-complex✕ fail9m 45s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps before the blue circle along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong shape

    agent's debrief

    The 8x8 grid with ~15 crossing arrows was far too tangled to trace by eye, so I did it computationally: extracted every black line component, detected the arrowhead blobs (local density spikes) and matched each head to its target shape and its line's tail shape. The arrow into the blue circle comes from the red diamond, and the arrow into the red diamond comes from the green diamond — so two steps before the blue circle is the green diamond. I was initially quite unsure reading this visually; the programmatic cross-check is what gave me confidence.

  • chart-simple✓ pass3m 11s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Bar chart with a bold main title at top; read directly.

  • chart-medium✓ pass47s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what value did Feb have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Bar chart 'New Signups'; the Feb bar tops out just above the 90 gridline, about 92.

  • chart-complex✓ pass7s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what is the difference between New and Returning in Feb? Answers within +/-4 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Website Sessions grouped bars, in thousands. In Feb: New (blue) is just under the 50 gridline (~48); Returning (orange) is about 35. Difference ~13.

  • screenshot-simple✓ pass6s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    AboStore cart panel: Total shown as $244.21; line totals 48.67+125.94+69.60 match.

  • screenshot-medium✓ pass7s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    AboStore cart panel: Total shown as $355.99; line totals sum matches.

  • screenshot-complex✕ fail12s

    prompt

    Look at the image at (fetch it and view it). What is the discount amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 0

    expected $38.96, got "-\$38.96"

    agent's debrief

    Order summary: Subtotal 649.36, Discount -38.96, Shipping 16.59, Tax 42.73, Total 669.72 (arithmetically consistent). Discount amount displayed is -$38.96.

  • diagram-simple✓ pass6s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Island" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Quiver -> Island -> Saddle -> {Basalt, Falcon, Cedar}. The arrow from Island points to Saddle.

  • diagram-medium✓ pass8m 46s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Turnip" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The bottom-row lines cross each other, so I traced computationally: walked up from the Guitar arrowhead at x=387 and it reached (503,318) right under Turnip (x=495); a second upward walk from the adjacent Guitar arrowhead (358) reached the Osprey area, and the Weasel arrowhead traced back to Osprey's right exit. A 4x zoom of the crossing region confirmed Turnip's line slants left, crosses Osprey's right line, and lands on Guitar. So Turnip -> Guitar.

  • diagram-complex✓ pass37m 33s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Birch"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Birch has exactly one incoming arrowhead, on its top border at x~1090. Connected-component + row-wise inertia trace: the arrowhead links to the stub at x=801,y=631 which sits on Pepper's bottom border (box x 728-847); the line runs down-right, crosses Jackal's lines near (888,657) and ends at Birch. Jackal's two bottom stubs (x=934,960) go to Turnip and Marmot respectively; the other top-border contact (x~1113) is a plain line from Juniper with no arrowhead.

Finding and reading email test

not examined · 0/6 answered

  • aggregate-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the inbox folder? Answer with just the number.
  • aggregate-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the archive folder have attachments? Answer with just the number.
  • temporal-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.
  • temporal-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the sent folder? Answer with just the subject line, exactly as shown.
  • needle-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message to gthorse@keyad.com about the Regatta, Sea Breeze & Harvard Place Apartments delivery, what is the airbill number given for the overnight shipment? Answer with just the number.
  • needle-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to gthorse@keyad.com about the Colonial Oaks property, what is the actual NOI for 2001 given for the property? Answer with just the number.

Purchasing test

not examined · 0/4 answered

  • find-product-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Automotive** category priced at or above **$950** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • find-product-2— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Toys & Games** category priced at or above **$100** with a rating of at least **3.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • purchase-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics - Luxury Side less Universal Fit Faux Leather Seat Cover Set with Steering Wheel Cover and Seat Belt Pads, Black (product id amazon.de:B07X33YLD7, abostore.airbench.ai/product/amazonbasics-luxury-side…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-a2fa1b7f@aidoctor.test. Answer with just the resulting order id.
  • recover-decline-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics 8 Sheet Strip Cut Shredder with CD Shred (product id amazon.com:B01E3R7GWA, abostore.airbench.ai/product/amazonbasics-8-sheet-str…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-7f3a4268@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

Coding test

not examined · 0/11 answered

  • compute-hash-1— unanswered—

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2742091716, 4080708501, 583986138, 89604131, 1609101632, 3845878337, 4260918646, 162940271, 1676292604, 2572099885, 2958005330, 914272507], x = 358948344, y = 3191915609 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.
  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 479 1: set b 818 2: set c 391 3: set d 484 4: mul a 78 5: mul a 56 6: mul a 14 7: dec d 8: jnz d -4 9: sub b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S....###...#.#..#...#...# .#...#.......#.....#.#... .#.#..#.##....#.......... .#..##.#......##..##.#.#. #.......#.....##..#.....# ..#..#.......###.##..#..# ###...#..#.#.#...##...... .##........#........#...# ##.#..##............#...# ...#.#.#.#....#.##..#...# ...#.....##.##........... #....##.#.......#...###.# .#......#.#...#..#..#...# ###.#...###....#.#..#.... ..#...##.......#..#.....# ..#...#...#.....###...### ...####...#..#..#..##.##. #.#.#..##.##......####..# #......#..#.#..#......#.. #.#...#.....###.#...#.... #..#..#....#...##.#.#...# .###........#.#.##....#.. .#..##.#..#.......##.#.#. .........#.#..#.......... ..#..##.#............#.#E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##.###.#..#.......#. ...#.##..##..##.#... ....#............... #...###.#...#.###.#. #.......#.#.#..##... .....#....#...##.... .#...#....#..#....#. ........#...#..###.. ......##.##.##.##.## .#....##.#.##..##... .#.......#.##....#.. ##..#....##.##...... ...##..###..#..#...# .#.....#....#.##.... ..#..#.#........##.# #...#.#....######... #.....#............. ...##..#.....#...##. #...#......#....#.#. ....##..#.#.#.#..... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 7930207892134786 and m = 15485863. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. Shatru bassha ficdor Pelfic lulu kaqui pelka, lubas ficdor? shatru, monix quivo lubas "tinix" ficdor Voti baszan vofic; ficdor basvo trulu "monix" basmo monix dorpel vofic lubas shamo, basvo Monix Monix basbas Lulu quika Karen shatru ficti! ficren; kaqui quivo "ficdor" Monix; Luka shamo quivo lubas basbas quivo shafic trulu tinix shamo ficti luka Tinix quivo lubas ficren basbas monix dorpel Voti quivo? Vofic Lulu Lubas! Kaqui monix mozan lubas shatru lubas lulu monix ficdor Nixfic ficdor ficdor quika Ficren Lubas mozan "Pelfic" lubas; basbas ficti nixfic shamo, nixfic shamo dorpel. basvo ficti Shamo ficti baszan shatru Nixsha basbas quivo bassha basvo basmo dorpel basbas basbas VOTI mozan nixzan Lubas ficren Ficti? luka tinix ficdor ficdor Shatru Dorpel pelfic lubas "shatru" ficdor vofic quika NIXSHA tinix shafic pelfic Lubas pelfic NIXFIC trulu quivo quivo baszan ficdor trulu Quivo monix trulu ficdor tinix ficren quivo tinix luka lubas Nixzan pelfic Pelka tinix basbas "lubas" Basmo lubas Shatru quivo. basvo ficdor ficti ficdor nixzan Kaqui? vofic pelfic basbas trulu basbas ficren dorpel lubas shamo mozan pelka ficti basbas pelfic quivo Dorpel lubas "tinix" Mozan pelfic baszan nixzan basvo basbas Shatru "kaqui" vofic! "pelfic" monix lubas basvo "lubas" lubas karen VOTI vofic. mozan Tinix nixzan dorpel tinix lulu "tinix" lulu, basvo trulu pelka; lubas? basvo ficdor "kaqui" ficdor shafic mozan kaqui lubas, mozan ficdor quivo shatru dorpel "ficren" ficti baszan shamo tinix. "Vofic" Basbas shamo luka "lubas" basvo monix Quivo shatru Baszan karen Nixzan dorpel ficdor baszan? ficren tinix monix "mozan" Nixzan Nixzan; shatru pelfic kaqui NIXZAN pelfic baszan basmo "Bassha" "ficdor" karen? bassha ficdor, pelka, Ficdor tinix monix Basvo tinix bassha Ficdor lubas nixsha Monix Tinix trulu luka mozan? quika. basvo trulu? dorpel. basbas Basvo dorpel ficti shamo! lubas, tinix Tinix tinix. tinix "bassha" Voti ficdor shafic monix karen monix monix mozan Luka pelka; pelfic Monix Shamo quivo shatru shatru quika quika lubas baszan? Lubas lubas dorpel ficti lubas lubas lubas lubas; voti pelka, Baszan kaqui basbas lubas basbas; quika lubas basvo monix lulu. pelfic Lubas tinix pelfic! ficdor Lubas; "tinix" shatru quivo Ficren mozan quika kaqui basvo bassha Monix ficdor vofic quika shatru lubas tinix nixsha tinix mozan trulu Quivo lulu nixzan! BASZAN shatru basvo shafic baszan pelka Basmo lubas ficdor shatru basbas monix Voti shatru dorpel Bassha mozan nixsha ficti trulu Lubas ficren mozan lubas basvo mozan nixsha voti lulu quika vofic pelfic nixzan? monix ficren basbas voti? shatru ficren; tinix Basbas Luka karen Dorpel, monix "tinix" "dorpel" Pelka lubas mozan quivo shamo mozan Quika Dorpel baszan Basvo, SHAMO lulu
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1fns = []; for (var v1i = 0; v1i < 2; v1i++) v1fns.push(() => v1i * 4); let v1 = 0; for (const f of v1fns) v1 += f(); const v2 = (0.1 * 5 + 0.2 * 5 === 0.3 * 5) ? "equal" : "different"; const v3 = ["6", "44", "101"].map(parseInt).join(","); const v4 = ["1" == 1, [] == false, NaN === NaN].map(Number).join(""); console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1591 cents, but the correct quote is 1971: {"country":"US","items":[{"grams":304,"qty":3,"price":1291,"fragile":true}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 482, 869, 1345, 1807]; // cents, by zone const PER_STEP = [0, 84, 133, 198, 251]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4100, 10000, 18500, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += 1; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":1271,"qty":3,"price":6320,"fragile":false}]} {"country":"BR","items":[{"grams":1387,"qty":4,"price":8898,"fragile":false},{"grams":290,"qty":2,"price":3021,"fragile":false},{"grams":308,"qty":4,"price":8660,"fragile":false},{"grams":247,"qty":2,"price":1804,"fragile":true}]} {"country":"FR","items":[{"grams":1395,"qty":5,"price":386,"fragile":false},{"grams":1404,"qty":4,"price":7674,"fragile":false},{"grams":344,"qty":1,"price":7641,"fragile":false}],"coupon":"SHIP10"} {"country":"JP","items":[{"grams":770,"qty":2,"price":5016,"fragile":false},{"grams":329,"qty":4,"price":706,"fragile":true}]} {"country":"MX","items":[{"grams":1201,"qty":2,"price":1985,"fragile":false},{"grams":304,"qty":5,"price":629,"fragile":false},{"grams":1004,"qty":3,"price":7716,"fragile":false},{"grams":974,"qty":4,"price":8785,"fragile":false}]} {"country":"US","items":[{"grams":263,"qty":5,"price":5155,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":177,"qty":2,"price":1013,"fragile":true}]} {"country":"FR","items":[{"grams":566,"qty":1,"price":3297,"fragile":false},{"grams":1769,"qty":3,"price":3509,"fragile":true},{"grams":1664,"qty":5,"price":1229,"fragile":false},{"grams":1202,"qty":5,"price":7128,"fragile":false}]} {"country":"MX","items":[{"grams":118,"qty":4,"price":3309,"fragile":false},{"grams":1781,"qty":5,"price":2640,"fragile":true},{"grams":1410,"qty":5,"price":4447,"fragile":false},{"grams":327,"qty":4,"price":971,"fragile":false}]} {"country":"ES","items":[{"grams":142,"qty":2,"price":2815,"fragile":true}]} {"country":"ES","items":[{"grams":114,"qty":2,"price":2347,"fragile":true}]} {"country":"BR","items":[{"grams":1465,"qty":5,"price":1782,"fragile":false},{"grams":420,"qty":2,"price":5386,"fragile":false},{"grams":1571,"qty":1,"price":1919,"fragile":true},{"grams":1089,"qty":1,"price":4720,"fragile":true}],"express":true} {"country":"US","items":[{"grams":416,"qty":3,"price":2753,"fragile":true}]} {"country":"DE","items":[{"grams":584,"qty":2,"price":773,"fragile":true}]} {"country":"ES","items":[{"grams":1471,"qty":1,"price":4140,"fragile":true}]} {"country":"AU","items":[{"grams":185,"qty":2,"price":1586,"fragile":true}]} {"country":"JP","items":[{"grams":918,"qty":5,"price":5717,"fragile":false},{"grams":1499,"qty":3,"price":3641,"fragile":true},{"grams":627,"qty":5,"price":7124,"fragile":false},{"grams":559,"qty":4,"price":6473,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":479,"qty":2,"price":876,"fragile":true}]} {"country":"BR","items":[{"grams":1144,"qty":1,"price":2971,"fragile":true}],"express":true} {"country":"ZA","items":[{"grams":451,"qty":2,"price":6545,"fragile":false},{"grams":1087,"qty":3,"price":1225,"fragile":false},{"grams":977,"qty":4,"price":453,"fragile":false}],"express":true,"coupon":"SHIP10"}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[12,19],[37,39],[40,40]] [[0,6],[35,40],[18,21]] [[31,38],[24,31],[40,41],[18,20],[28,35],[9,11],[38,43]] [[21,23],[40,44],[9,16],[4,6],[6,9]] [[33,38],[32,40],[40,45],[24,30],[11,15]] [[26,32],[24,25],[19,26],[24,30],[21,23],[25,29],[7,9],[20,23]] [[25,25],[14,14],[27,33],[15,16],[23,24],[25,28]] [[12,16],[16,18],[10,14]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/8462b03e52ad08611a45ed1d900f3bd8.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/28211cb4d77950ea81375c678895a0a5.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 (quantization-aware-trained NVFP4, compressed-tensors, MTP head kept). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95 --speculative-config '{"method":"mtp","num_speculative_tokens":2}'. Harness: openclaw 2026.9.6 in a container (node:24): `openclaw agent exec --config <pinned per-run file> --state-dir <workspace> --json <prompt>`; provider api openai-completions; context 131072, max output 32768 tokens; agents.defaults.compaction.midTurnPrecheck.enabled=true, compaction.timeoutSeconds=1800; everything else openclaw's exec defaults. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 2a84999; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.

conclusion

Result: 26 passed, 2 failed, 21 unanswered (math 9/9, vision 17/19; mail, purchasing and coding not reached). Wall time 119 min. openclaw compacted the ~111k-token session successfully, but its compaction safeguard reported 'finalized artifact truncated; loss=summary-tail': the summary was cut to fit openclaw's summary budget (the room left after the recent messages it keeps verbatim), dropping its end. Three tool calls later the model ended the run (stopReason=stop) without starting the remaining sections, apparently because the outstanding work list was in the lost tail. An openclaw compaction behaviour, not a model or server failure. Written by the orchestrator (agent-checkup-benchmark) from the run's logs and diagnosis.

discussion

Sign in to join the discussion

No messages yet.