airbench.ai

Benchmark v1.0 · report

openclaw/rtx5090/qwen3.8-27b-nvfp4

sharedairbench.ai/checkup/b90a597f-d6c0-4b18-80ab-b3fb6636582d/report

setup

model type
open model (local)
hardware
RTX5090
harness
openclaw
model
qwen3.8-27b-nvfp4
modelself-reportedhal/qwen38-27b-nvfp4-128k

started 2026-09-26 11:53 UTC · shared 2026-09-27 16:45 UTC

overall

Answered 26 of 49 challenges; 24 correct.

24 of 49 challenges passed

partial run · 23 unanswered, counted against the score

  • 24 passed
  • 2 failed
  • 23 not answered

vitals

time

24m 30s

answered

53%

failed

4%

success

49%

systems

Math test

8/9 passed

time to last answer 1m 53s
  • letter-count-1✓ pass57s

    prompt

    How many times does the letter "d" appear in "ddordkamddo"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial. Counted d,d,o,r,d,k,a,m,d,d,o - five d at positions 0,1,4,8,9. No difficulty.

  • decimal-compare-1✓ pass4s

    prompt

    Which decimal number is larger, 3.2 or 3.21? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial: 3.21 is larger than 3.2 (3.2 = 3.20). No tricks here.

  • arithmetic-1✓ pass8s

    prompt

    Compute step by step, left to right (no operator precedence): 25 - 2 * 2 / 2 * 3. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward left-to-right: 25-2=23, *2=46, /2=23, *3=69. Verified with a script to avoid arithmetic slips.

  • unit-convert-1✕ fail14s

    prompt

    Convert 19 GB to MB. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Chained conversions: 19 GB * 1024 MB/GB = 19456 MB, then 19456 hours * 60 = 1167360 minutes. I used 1024 MB per GB (the conventional computer-memory convention); the prompt does not say whether SI (1000) was intended, which would give 1140000. Ambiguity is on the challenge, not the arithmetic.

  • format-json-1✓ pass13s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "9439". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 9439. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The prompt demanded an exact-ordered JSON object with a numeric checksum. Straightforward: digit sum of 9439 is 25. I generated the inner JSON with json.dumps to guarantee valid syntax and key order. My only real worry was keeping the outer submission body valid while embedding JSON inside JSON, which I handled by writing the body with a script.

  • math-add-1✓ pass16s

    prompt

    What is 4 + 10? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition: 4+10=14. No issue.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 555 + 584. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    555+584: 555+584 = 555+600-16 = 1139. Verified by complementing. Routine.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-10 + -6) * (22 - 29)) + (5 * 2) - 36

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    (-10 + -6) = -16; (22-29) = -7; product = 112; 5*2=10; 112+10-36 = 86. Verified by hand and re-check. Routine.

  • math-determinant-1✓ passbatched

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [10, -8, -1, 3] [5, 7, -5, -2] [6, 6, 6, -9] [8, 7, 1, 11]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    4x4 determinant. I computed it two independent ways - exact fraction-based Gaussian elimination and full permutation (Leibniz) expansion over 24 terms - and both gave 21489, so I am confident. No ambiguity.

Vision test

16/19 passed · 2 unanswered

time to last answer 24m 30s
  • acuity-20✓ pass3m 26s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Eye chart row 4 group 2. I read the full chart, then cropped and zoomed that exact region to confirm the characters - it shows 2DCUM cleanly. Confident.

  • acuity-14✓ pass46s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 3 is small text. I read the full chart, then located the row band programmatically and zoomed the crop - it shows N6D5U. One of my first crop attempts was mis-aimed (blank image), so I recomputed coordinates from a pixel scan of the actual row bands. Confident.

  • acuity-10✓ pass42s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 group 1 is tiny. My first zoom crop actually landed on row 7 (8WD39) because I mis-estimated the y coordinates; I corrected by re-cropping one band up. The final zoom shows WB94B clearly, matching my initial whole-chart read.

  • acuity-8✓ pass22s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 group 1, the smallest row on the chart. I read the full image, then zoomed that region to double-check. The zoom is pixelated at the edges but the code reads NZ49E. I am moderately confident - this is the limit of what the tiny font resolves to.

  • count-simple✓ pass29s

    prompt

    Look at the image at (fetch it and view it). How many orange squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted orange squares visually as 5, then verified with a connected-component flood fill on the exact orange color (242,106,34). Found exactly 5 identical-size components. The blue and purple diamonds and red triangle are different colors and correctly excluded.

  • count-medium✓ pass1m 01s

    prompt

    Look at the image at (fetch it and view it). How many teal diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    15 teal shapes total. Classified each by width profile: 12 diamonds (narrow top and bottom), 3 triangles (wide base). Confident in 12.

  • count-complex✓ pass39s

    prompt

    Look at the image at (fetch it and view it). How many green circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted 39 green circles by programmatic flood-fill on the exact green (22,163,74), classifying each component by shape signature: 39 circles (fill 0.80), plus 4 triangles, 3 squares, 2 diamonds correctly excluded. Teal excluded by the blue-channel threshold. All components uniform size, so no merges to worry about.

  • spatial-simple✓ pass15s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    5x5 grid, one red circle. Unambiguously at row 3, column 2 - it is the only red shape on the board, so no confusion with orange circles nearby.

  • spatial-medium✓ pass6m 12s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the orange circle lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tricky: I traced all 7 arrows programmatically (black-pixel components, arrowhead detection via local density, nearest-shape matching) and confirmed by zooming in. The arrow from the orange circle (row 4, col 6) goes up-left and its head points at the red triangle. My initial trace was contaminated by antialiased grid-line pixels misclassified as a fake shape, so I had to filter those out. Confident in red triangle.

  • spatial-complex✕ fail4m 15s

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the red diamond along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong count

    agent's debrief

    8x8 grid with 16 arrows - the hardest one so far. I traced all arrows programmatically (black-pixel components, arrowhead density, endpoint matching) and verified the red diamond has exactly one outgoing arrow by zooming. Chain: red diamond -> teal triangle -> red triangle = 2 shapes after it. First attempt crashed on a tuple-unpack bug and misread colors from quantized bins; fixed both before trusting the result.

  • chart-simple✓ pass29s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Title at the top of the bar chart. Read directly - it is the large bold text: New Signups. Routine.

  • chart-medium✓ pass57s

    prompt

    Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what is the difference in value between Mar and Jan? Answers within +/-8 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Measured instead of eyeballing: located gridlines (100/80/60 at y=119/227/335, 108px per 20 units, baseline 659), read bar tops (Jan 488, Mar 434). Jan ~31.7, Mar ~41.7, difference = 10. Well within the +/-8 tolerance.

  • chart-complex✓ pass2m 39s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, how many months did New have a value greater than 26? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted New-signup bars above 26 by measuring bar tops against the 100/75/50 gridlines (y=119/259/399) with 0-line at y=680 (confirmed by the dark axis). 11 of 12 months are >26; Dec is ~15.9 so it is below. My first pass had contaminated the baseline with the legend swatch and miscounted Dec as ~26.9 - caught that when Dec sat dangerously close to threshold and re-measured with the legend excluded.

  • screenshot-simple✓ pass32s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Cart panel total read directly as $161.41. Sanity-checked the arithmetic: 24.34 + 99.84 + 37.23 = 161.41 exactly. Routine.

  • screenshot-medium✓ pass33s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read the Total as $362.36. Verified line items: 174.44+106.92+24.00+57.00 = 362.36 exactly, and each line total matches qty*unit. Routine.

  • screenshot-complex✓ pass37s

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Order summary lists Subtotal 346.63, Discount -24.26, Shipping 19.87, Tax 22.57, Total 364.81. Asked for shipping specifically: $19.87. Cross-checked the whole panel arithmetic (346.63-24.26+19.87+22.57 = 364.81) - it is consistent, so I trust the printed figure.

  • diagram-simple✓ pass35s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Sequoia"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Small DAG: Anchor branches to Alder and Dingo; Alder points to Sequoia; Dingo points to Ridge. So the only arrow into Sequoia comes from Alder. Clear.

  • diagram-medium— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Canyon"? Answer with just the box name, e.g. Kettle.
  • diagram-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Orbit" point to? Answer with just the box name, e.g. Kettle.

Finding and reading email test

not examined · 0/6 answered

  • aggregate-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the inbox folder have attachments? Answer with just the number.
  • aggregate-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include jacquestc@aol.com in the To field? Answer with just the number.
  • temporal-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.
  • temporal-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.
  • needle-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the reminder about the Portland Fundamental Analysis Strategy Meeting, what participant code is given for the call-in? Answer with just the number.
  • needle-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Zero Option", what dollar amount is given for the outstanding bill that will hit Enron in Q1 2002? Answer with just the number.

Purchasing test

not examined · 0/4 answered

  • find-product-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Grocery & Gourmet** category priced at or above **$800** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • find-product-2— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Toys & Games** category priced under **$80**, which has the **highest rating**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • purchase-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Cut-End Cotton Commercial String Mop Head, 1.25 Inch Headband, Medium, Blue, 6-Pack (product id amazon.ca:B071X93RGW, abostore.airbench.ai/product/amazonbasics-cut-end-cot…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-45eb0e45@aidoctor.test. Answer with just the resulting order id.
  • recover-decline-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Basic Care Ibuprofen Tablets 200 Count (product id amazon.com.au:B074F18H5P, abostore.airbench.ai/product/basic-care-ibuprofen-tab…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-e7881254@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

Coding test

not examined · 0/11 answered

  • compute-hash-1— unanswered—

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2813209872, 2959750225, 2540143302, 2715969279, 3018261708, 503457341, 2357887138, 651296139, 2996295624, 3783171177, 2831703998, 1702082135], x = 3021554180, y = 884990677 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.
  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 843 1: set b 335 2: set c 353 3: set d 577 4: sub a 66 5: sub b a 6: add a 42 7: dec d 8: jnz d -4 9: add b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S....##...#.#....####...# ...........#....#........ ##....#....#....##....... .#.............#.....#... .........#.#.##....##.... .#.#.#....#.....#..#..... .............##..#.....#. .........#...#....#.#..#. #...#.#...#.....#.......# ..##...#.#...#........... #....##.............###.# #..#....#..#......##.#.#. ..#....#.#.#..##......#.. ..#.#.#..#............... ..#..#.#....#......#.#... .#....###.....#...#....#. ......#......#.#...#.#..# ..#...#....#.#..#...#.### ........#....#...#....#.. ##.#.#...#..##..#.....##. .........#.....#.#...#.## ..#....#.....##..###..... #........#.#....#........ .....#...#............... #..#.#...#..#.###..#...#E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: #........#.##.#....# ....#..#.....#.#.### ....#.......#.##.#.. #..#..###...#.#..### ...#.#.##....#..#.## #......#.....#..##.# .#...#..##.#..#....# #....#...#..#..#...# ..#.#...#.#......... ..#....####......#.# #.#....#......#..... .#.###..#.#...##.... .#.........#.#.##... .#...#...###.#...#.# .......#.###..#.#..# ...#.#..#...#.###... .###....#.....#....# ...#......#.###.###. #.#..#...#.....##.## ......##.##...###.## Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 1415364213776774 and m = 15485863. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. voti nixdor "quibas" votru, pelbas Zantru. "Luqui" ficmo lunix shaqui Nixbas Shaqui Quibas pelfic kador monix tipel BASBAS trutru mofic renvo nixti bastru! renmo Renvo vonix luqui; Luqui Kador Basbas vonix basbas vonix lunix voti nixdor kaka shaqui? zantru renmo renmo renvo mofic Moqui renvo kador zantru basbas renmo Monix nixdor renmo shaqui Vonix "truka" Molu, SHAQUI, votru? moqui. Votru nixti renmo Basbas? Voti ficmo vonix nixdor renmo mofic basbas luqui votru Renmo monix momo lunix Vonix? Luqui Nixbas moqui moqui, MOLU? Renvo Basbas basbas molu Basbas quibas "basbas" Kaka tiren bastru nixbas zanqui renvo molu ficmo kaka "vonix" votru renmo RENVO. "VOTI" votru renmo! Luqui "ficmo" Quibas tiren shaqui, vonix pelbas pelfic tipel Mofic Pelfic Zanqui zanqui BASBAS zantru renmo nixbas VOTRU! RENMO? kaka quibas bastru moqui KAKA. "renvo" bastru. basbas momo vonix truka zantru Pelfic shaqui renmo Momo? shaqui voti Nixti lunix pelfic, "luqui" basbas moqui! renvo zantru lunix trutru! Mofic lunix basbas "mofic" Monix RENMO votru "moqui" basbas tiren mofic shaqui tipel nixdor LUQUI zantru nixbas momo shaqui ZANQUI Molu renmo basbas renvo moqui "Zanqui" molu lunix, renmo luqui, nixdor renmo moqui, zantru momo trutru molu kador Lunix renvo Renmo pelfic nixti renvo kador trutru Pelfic molu renmo tiren renmo mofic Tipel Renmo "Basbas" Bastru zanqui molu PELFIC Renmo tiren. zanqui "tipel" renmo zantru pelfic mofic bastru shaqui renmo Pelfic? kador nixti truka nixdor trutru nixbas Renmo shaqui Zanqui moqui Tipel Momo! renmo votru tiren shaqui LUNIX. quibas "Kador" pelfic renmo molu Tiren lunix basbas zantru Basbas tipel truka molu! luqui renvo mofic tiren moqui basbas renvo Shaqui ficmo nixbas zantru tipel renmo pelbas NIXBAS nixdor Luqui zantru pelfic "shaqui" Zantru renmo Quibas moqui? renmo votru luqui molu basbas pelbas zantru. Nixbas tipel basbas mofic ficmo molu Ficmo Pelbas renmo Renvo votru Zantru molu basbas pelbas renmo nixbas shaqui nixbas? basbas nixbas zantru TRUKA kaka renmo renvo shaqui shaqui molu pelfic? "Renmo" Pelfic votru tiren Trutru shaqui ZANTRU renmo lunix Nixbas. Luqui zantru! Mofic monix trutru trutru votru molu, VONIX renmo ficmo renmo, nixbas luqui Basbas pelbas "ficmo" "zanqui" VONIX renmo momo. nixdor, Truka Renmo moqui? kador renmo voti! basbas quibas nixdor lunix, QUIBAS! bastru luqui zantru, Nixbas nixbas vonix zanqui tiren vonix, pelbas Bastru? momo mofic voti? pelfic renvo lunix "pelfic" renmo bastru trutru lunix shaqui? zanqui nixti molu luqui mofic Monix votru kaka Renmo Basbas renmo mofic basbas luqui shaqui trutru RENVO zantru kaka Kaka basbas Nixdor basbas renmo zanqui ficmo Tipel basbas nixbas! renmo ficmo nixbas tipel nixbas Pelfic, Lunix! Shaqui basbas mofic Basbas
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = ["5", "21", "110"].map(parseInt).join(","); const v2 = "9" + 4 - 2 + "2"; const v3 = [38 / 7 | 0, Math.round(-8.5), -96 % 8].join(","); const v4 = [null >= 0, NaN === NaN, null == 0].map(Number).join(""); console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 156 cents, but the correct quote is 234: {"country":"IT","items":[{"grams":359,"qty":2,"price":2849,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 507, 835, 1235, 1735]; // cents, by zone const PER_STEP = [0, 78, 143, 226, 262]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5600, 9000, 18100, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"CA","items":[{"grams":378,"qty":5,"price":1107,"fragile":false}]} {"country":"BR","items":[{"grams":587,"qty":3,"price":4724,"fragile":true}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":351,"qty":3,"price":7433,"fragile":false},{"grams":986,"qty":4,"price":6476,"fragile":false},{"grams":95,"qty":1,"price":8016,"fragile":false}]} {"country":"FR","items":[{"grams":660,"qty":5,"price":2267,"fragile":false}]} {"country":"IT","items":[{"grams":843,"qty":4,"price":308,"fragile":false}]} {"country":"DE","items":[{"grams":1151,"qty":3,"price":7779,"fragile":false},{"grams":1716,"qty":1,"price":5040,"fragile":false}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":1195,"qty":1,"price":8536,"fragile":false},{"grams":1540,"qty":1,"price":8644,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":1297,"qty":1,"price":590,"fragile":false}]} {"country":"CA","items":[{"grams":335,"qty":2,"price":679,"fragile":false}]} {"country":"MX","items":[{"grams":1422,"qty":2,"price":7480,"fragile":true},{"grams":892,"qty":4,"price":4520,"fragile":false},{"grams":383,"qty":1,"price":5909,"fragile":true}]} {"country":"DE","items":[{"grams":304,"qty":2,"price":1069,"fragile":false}]} {"country":"FR","items":[{"grams":825,"qty":5,"price":2100,"fragile":false}]} {"country":"IT","items":[{"grams":716,"qty":5,"price":8606,"fragile":false},{"grams":255,"qty":2,"price":3240,"fragile":false},{"grams":795,"qty":2,"price":1339,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":1363,"qty":5,"price":7835,"fragile":false},{"grams":841,"qty":3,"price":3128,"fragile":false}]} {"country":"FR","items":[{"grams":514,"qty":4,"price":2423,"fragile":false}]} {"country":"GB","items":[{"grams":1493,"qty":3,"price":3997,"fragile":true},{"grams":982,"qty":1,"price":6608,"fragile":true}],"express":true} {"country":"CA","items":[{"grams":1017,"qty":4,"price":1389,"fragile":true}]} {"country":"CA","items":[{"grams":844,"qty":4,"price":7623,"fragile":false},{"grams":1329,"qty":5,"price":8988,"fragile":false},{"grams":304,"qty":3,"price":7504,"fragile":false}],"coupon":"SHIP10"} {"country":"DE","items":[{"grams":564,"qty":2,"price":8850,"fragile":false},{"grams":1245,"qty":3,"price":6838,"fragile":false},{"grams":862,"qty":2,"price":5751,"fragile":true},{"grams":1772,"qty":2,"price":1211,"fragile":false}]} {"country":"DE","items":[{"grams":1555,"qty":3,"price":3475,"fragile":true}],"coupon":"SHIP10"}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[40,44],[23,24],[15,16]] [[25,32],[25,27],[17,23],[4,9]] [[25,32],[13,19],[1,4],[33,33],[18,26],[32,32],[30,37]] [[14,14],[16,24],[5,12],[12,20],[30,31],[23,27]] [[3,9],[0,7],[29,35],[24,25],[4,6],[20,27]] [[24,24],[20,22],[8,8],[22,27],[30,34],[36,43],[25,31]] [[14,21],[9,14],[39,40],[29,33],[1,4],[6,13]] [[4,4],[13,17],[9,16],[3,9]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/70958b0d30f41f0d89cf325054c56bdd.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/8066fdeb0e75323ba68df478d6801982.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 (NVFP4, no MTP head). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --quantization modelopt --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95. ~79 tok/s single-stream decode. Harness: openclaw 2026.9.6 in a container (node:24): `openclaw agent exec --config <pinned per-run file> --state-dir <workspace> --json <prompt>`; provider api openai-completions; context 131072, max output 16384 tokens; everything else openclaw's exec defaults. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 6d738a5; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted.