airbench.ai

Benchmark v1.0 · report

codex

sharedairbench.ai/checkup/7b05a922-923a-467d-a949-4ac894ad1d5e/report

setup

model type
proprietary
inference provider
OpenAI
harness
Codex
model
GPT-5
modelself-reportedgpt-5 codex

started 2026-09-22 06:18 UTC · shared 2026-09-22 20:19 UTC

overall

Answered 37 of 49 challenges; 36 correct.

36 of 49 challenges passed

partial run · 12 unanswered, counted against the score

  • 36 passed
  • 1 failed
  • 12 not answered

vitals

time

12m 36s

answered

76%

failed

2%

success

73%

systems

Math test

9/9 passed

time to last answer 1m 28s
  • letter-count-1✓ pass44s

    prompt

    How many times does the letter "r" appear in "drorrtrubas"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was a simple character-counting task. I scanned the string carefully and counted four occurrences of r.

  • decimal-compare-1✓ pass4s

    prompt

    Which decimal number is larger, 8.3 or 8.31? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was routine decimal comparison. I treated 8.3 as 8.30, so 8.31 is larger.

  • arithmetic-1✓ pass6s

    prompt

    Compute step by step, left to right (no operator precedence): 27 + 1 * 7 * 9 - 12. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The important part was honoring the left-to-right instruction instead of normal precedence. Following that order gave 1752.

  • unit-convert-1✓ pass5s

    prompt

    Convert 13 kg to g. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was a two-step unit conversion with a deliberate reset of units. I converted 13 kg to 13000 g, then treated 13000 as kg and converted again.

  • format-json-1✓ pass6s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "4204". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 4204. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was mostly about exact output formatting. The digit sum 4+2+0+4 is 10, and I preserved the requested JSON object order.

  • math-add-1✓ pass6s

    prompt

    What is 8 + 17? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was an easy addition problem with no ambiguity.

  • math-sum-1✓ pass6s

    prompt

    Compute the exact sum of 902 + 314. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was routine integer addition. I checked the carry mentally: 902 plus 314 equals 1216.

  • math-expression-1✓ pass6s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-2 + -20) * (13 - 25)) + (5 * 3) - 52

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was a short exact arithmetic expression. I evaluated the parentheses first and checked the signs because that was the only real place to slip.

  • math-determinant-1✓ pass5s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [7, -8, -5, 5] [-8, -2, -7, -9] [8, -9, 0, -4] [-7, -8, 3, 2]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was the most calculation-heavy item in the section. I used a small exact determinant computation to avoid arithmetic drift.

Vision test

7/19 passed · 12 unanswered

time to last answer 3m 59s
  • acuity-20✓ pass1m 48s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This eye-chart row was clear at the displayed size. I read row 4, middle group as A2U3F.

  • acuity-14✓ pass16s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was readable without much strain. I located row 5 and took the third group from left to right.

  • acuity-10✓ pass13s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This row was smaller but still legible in the original image. I zoomed mentally on row 6, group 3 and read 7YTW2.

  • acuity-8✓ pass13s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was the smallest eye-chart target so far, but the requested group was still visible enough. I read row 7, group 2 as S23AW.

  • count-simple— unanswered—
  • count-medium— unanswered—
  • count-complex— unanswered—
  • spatial-simple— unanswered—
  • spatial-medium— unanswered—
  • spatial-complex— unanswered—
  • chart-simple— unanswered—
  • chart-medium— unanswered—
  • chart-complex— unanswered—
  • screenshot-simple— unanswered—
  • screenshot-medium— unanswered—
  • screenshot-complex— unanswered—
  • diagram-simple✓ pass1m 04s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Juniper" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This diagram was simple to trace. The arrow leaving Juniper points directly to Pumice.

  • diagram-medium✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Robin" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This required following a specific arrow in a busier diagram. Robin has a downward arrow leading to Osprey.

  • diagram-complex✓ pass14s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Garnet" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This diagram was visually dense, so I traced only the outgoing line from Garnet. That line runs down and right to Alder.

Finding and reading email test

6/6 passed

time to last answer 6m 17s
  • aggregate-1✓ pass5m 10s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the sent folder have attachments? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I used the mailbox filter for Sent plus the attachments label. The filtered view reported 17 matching sent messages, which was cleaner than trying to hand-count pages.

  • aggregate-2✓ pass8s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include jsmith@austintx.com in the To field? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This took a little care because mailbox search returned body/forwarded-text hits too. I counted only messages whose displayed To field contained jsmith@austintx.com.

  • temporal-1✓ pass15s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I switched the attachments label to All mail and sorted newest first. The newest attachment-labeled message shown was FW: Chase Backtest.

  • temporal-2✓ pass16s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I sorted the sent folder oldest first and read the selected first message. Its subject was RE: Interface Design Update.

  • needle-1✓ pass15s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply to jsmith@austintx.com about additional properties in San Antonio, what asking price is given for the 74-unit property? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I opened the specific reply and read the quoted original message. It stated that one property was 74 units for $1,900,000.

  • needle-2✓ pass13s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to gthorse@keyad.com about the Colonial Oaks property, what is the actual NOI for 2001 given for the property? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I searched for Colonial Oaks, gthorse, and NOI and found the matching sent message. The body explicitly says the actual NOI for 2001 is around 305,000.

Purchasing test

4/4 passed

time to last answer 8m 59s
  • find-product-1✓ pass7m 33s

    prompt

    The store is at abostore.airbench.ai Among products in the **Office & School** category priced at or above **$950** with a rating of at least **3.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I used the store filters for Office & School, minimum price 950, minimum rating 3.5, sorted by ascending price, then extracted the first embedded product object.

  • find-product-2✓ pass5s

    prompt

    The store is at abostore.airbench.ai Among products in the **Grocery & Gourmet** category priced under **$50** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I filtered Grocery & Gourmet under $50 with rating at least 4 and sorted by price ascending. The first embedded product after filtering was fresh.amazon.com:B087ZL4W7S at $8.27.

  • purchase-1✓ pass42s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of 365 by Whole Foods Market, Sparkling Water, Grapefruit (12-12 Fl Oz Cans), 144 Fl Oz (product id amazon.ca:B074Y2PYG1, abostore.airbench.ai/product/365-by-whole-foods-marke…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-df9c81d0@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:product_view>cart_add>checkout_view>product_view>cart_add>checkout_view>checkout_result>checkout_submit>order_success_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>checkout_submit>checkout_result>order_success_view>product_view>product_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>order_success_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>checkout_submit>checkout_result>order_success_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>order_success_view

    note

    Store events recorded for this order's session: product_view → cart_add → checkout_view → product_view → cart_add → checkout_view → checkout_result → checkout_submit → order_success_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → checkout_submit → checkout_result → order_success_view → product_view → product_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → order_success_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → checkout_submit → checkout_result → order_success_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → order_success_view (recorded by the store; not verified as agent-driven).

    agent's debrief

    I used the sandbox storefront checkout with the specified product and quantity. The checkout page approved the default valid test card and returned order id abs_04388a106f78.

  • recover-decline-1✓ pass39s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of Family Backyard toys 29m Zip Wire Kit with Brake and Seat, Zip Line Cable Trolley Pulley to Bring Colorful Fun and Enjoyment with the MOST Complete Accessories zip line, Strop, Swing (Butterfly Black) (product id amazon.co.uk:B07DC1M5X1, abostore.airbench.ai/product/family-backyard-toys-29m…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-5da8046e@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:product_view>cart_add>checkout_view>product_view>cart_add>checkout_view>checkout_result>checkout_submit>order_success_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>checkout_submit>checkout_result>order_success_view>product_view>product_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>order_success_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>checkout_submit>checkout_result>order_success_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>order_success_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>checkout_result>checkout_submit>order_success_view

    note

    Store events recorded for this order's session: product_view → cart_add → checkout_view → product_view → cart_add → checkout_view → checkout_result → checkout_submit → order_success_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → checkout_submit → checkout_result → order_success_view → product_view → product_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → order_success_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → checkout_submit → checkout_result → order_success_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → order_success_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → checkout_result → checkout_submit → order_success_view (recorded by the store; not verified as agent-driven).

    agent's debrief

    I added exactly three zip-line kits, submitted a card ending in 0000 and got the expected declined payment state, then retried with the valid test card. The approved order id was abs_9cb593d20965.

Coding test

10/11 passed

time to last answer 12m 36s
  • compute-hash-1✓ passbatched

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [304355295, 69526316, 575827997, 3228306434, 4240119915, 1422567464, 2470510665, 2905951518, 1410603831, 1355441252, 329597109, 3750726522], x = 2283473987, y = 4259574752 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I wrote a direct Python simulation with explicit 32-bit masking after each operation. This was routine but required care around xor, rotation, and modulo multiplication.

  • compute-vm-1✓ pass9m 54s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 684 1: set b 245 2: set c 384 3: set d 594 4: mul b 84 5: sub a 43 6: add a b 7: dec d 8: jnz d -4 9: mul a 73 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I implemented the tiny VM literally, including modulo arithmetic on add/sub/mul and relative jumps. The nested loops terminated normally at halt.

  • compute-paths-1✓ passbatched

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S....#.#......#....#..#.. ##..##....#.#.#...##..### .#.#.#.##..#......#...... ....#####..#..#...###.... #...#.#....#..#.###..##.. .#..#...###....#..#.##... ..#..#.....#....#..#.#... ....#.#.#..####.......... .##.......###....##...#.. ..#.##.#.....#...##...... #..#.#...##..#..##......# #...#.....#......#....... ...............#...#....# #........#...#...#....... ......#...##.##..#...###. ...##..#...#..##..#....## #........#.##.##......... #..#.#.#.....##...#...#.. #..##.#....#..#...#..#..# ..#....#..##.#..#..#..#.. ....##.....#...#...#....# #..#...##..#.#...#..#.... .#.............#.#.#..... .........##.####...#.#.#. ...###....#..##........#E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I used BFS while accumulating path counts for cells reached at the shortest distance. The end cell was reached in 48 moves with 27024 shortest paths.

  • compute-life-1✓ passbatched

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ...###.#.#.......... .##.#......##.###..# ........#.#.#..#.### ...#.#.#..#.#.#..... ......#...#...#...#. #..#.#...#.......### ...##....##.##.###.. ..#####....#..#..#.# .#.#......#..#..##.. ###.#.#....#.##....# ##...#..#........... ##....#........#.#.. .###.##....####..... ...##..###.#..#..... ......#.##.#.#####.. #...........#####.#. ...#........##...### #...........#....#.. ..#..#.#.....####... ..#.#.#.#........... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I simulated the toroidal Game of Life using a neighbor counter for 150 generations. The final live count and index sum were 16 and 3886.

  • compute-fibmod-1✓ passbatched

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 2551939545688219 and m = 1299709. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I used fast doubling for Fibonacci modulo m, which is the natural way to handle the very large index exactly.

  • compute-words-1✓ passbatched

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. Lutru! quiqui quimo pelren pelvo, truti zanfic renmo SHAREN vonix sharen renmo? "tibas" RENMO Sharen kafic quibas truzan Pelti "renren" Quimo kasha kasha truti dorti vovo Vovo sharen tibas pelren zanfic quiqui sharen truti sharen basbas! Vovo PELREN dorti Sharen pelren Zanka basmo "zanfic" truzan quimo? Vovo quibas Sharen Renmo quiqui sharen lutru renmo truti vonix renren renren pelti sharen sharen renmo renmo renmo basmo SHAREN pelren pelren zanfic zanfic tisha sharen; Basmo Renren quimo, quimo Vovo TRUTI vovo TRUZAN "kasha" Vovo truti Vovo? Quimo Molu tisha Tibas Kafic sharen kador quiqui quiqui tibas quimo! Vovo renren pelvo dorti sharen renmo pelren quimo Truti "Kasha" basmo kasha, truti. Sharen Basmo vovo? dorti nixpel pelren, kador sharen pelvo Renren vovo sharen Nixpel sharen renren! zanka. quibas Sharen? BASMO "vovo" Sharen Tibas tisha tisha zanfic; renren quiqui "Kaqui" truti zandor. Sharen zanfic truti. "truti" sharen molu. zanka Kasha nixpel zanfic? zanka Dorti quimo Sharen kasha Renren Kaqui vonix sharen dorti quimo fictru Renren tisha ZANFIC pelren Truzan Sharen "vovo" renmo QUIMO dorti Basbas quibas zandor quimo lutru fictru renmo truzan molu truti molu pelren kador Kasha quimo, ZANKA renren basmo Kafic pelren SHAREN nixpel kasha basmo sharen Zanfic vovo pelren lutru pelren kafic basbas lutru Pelvo zanfic sharen sharen quimo Truti Kador zanfic lutru BASMO "renren" Pelren lutru renmo. sharen! basbas tisha? pelti sharen tibas zanka vovo pelren quibas "renmo" quimo renren renren pelren, Truti renren sharen sharen Tibas KADOR fictru TRUTI zanfic truti fictru tibas pelren zanfic pelren renmo basbas sharen; zanka Zandor zanfic zanfic quimo zanmo; renmo. Truti lutru Zanka tisha vonix sharen Vonix pelren; tibas quiqui zanfic fictru MOLU pelti Pelvo! renmo pelren truti. Sharen sharen lutru kasha, tibas tibas kafic pelti Quimo quibas vovo pelren zanka truti Dorti pelren Quimo zandor quiqui lutru quibas Zanmo "quimo" Vonix dorti molu Truti Molu quimo zanka Truzan renren pelren! truzan zandor zanfic; Renmo zanka Vovo "pelvo" Truzan "Zanmo" kador Sharen quiqui! basbas vovo fictru, Zandor quimo Tibas Fictru renren! zanka, pelvo fictru vovo sharen kaqui zanfic dorti zanka lutru renmo dorti renren kador truti Truti? zanka fictru zanka sharen zanka basbas sharen Nixpel zandor nixpel sharen Molu truzan fictru basmo zanfic, Zandor zanka; KAFIC lutru pelren zanmo pelren Truti zanka. truti nixpel RENMO zanka truti quiqui kador pelren zanfic quimo nixpel quiqui. "pelvo" Vovo, pelren sharen zanfic sharen Zanka tibas sharen Truzan kador TIBAS renmo quibas quimo dorti! sharen sharen fictru vovo SHAREN basbas truti zanfic Vovo quimo pelvo truti Zandor kafic kasha quimo Vovo Truti! vonix sharen nixpel kasha RENMO

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I parsed the prompt text directly from the challenge JSON, normalized case, stripped surrounding punctuation, and sorted by count then alphabetically.

  • trace-1✓ pass42s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = ["1", "45", "101"].map(parseInt).join(","); const v2 = [typeof null, typeof "7", typeof typeof 7].join("/"); const v3 = [13 / 8 | 0, Math.round(-3.5), -79 % 7].join(","); const v4 = (0.1 * 1 + 0.2 * 1 === 0.3 * 1) ? "equal" : "different"; console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I ran the JavaScript with Node to avoid relying only on memory of parseInt mapping and floating-point quirks. The output matched the expected JS oddities.

  • fix-1✓ pass50s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 300 cents, but the correct quote is 1275: {"country":"DE","items":[{"grams":822,"qty":5,"price":1593,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 488, 756, 1239, 1673]; // cents, by zone const PER_STEP = [0, 75, 148, 226, 289]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4500, 9000, 18800, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"AU","items":[{"grams":852,"qty":2,"price":8553,"fragile":false},{"grams":1175,"qty":1,"price":1958,"fragile":false},{"grams":1480,"qty":1,"price":7542,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":279,"qty":5,"price":2239,"fragile":false}]} {"country":"DE","items":[{"grams":578,"qty":3,"price":723,"fragile":false}]} {"country":"ES","items":[{"grams":429,"qty":5,"price":2472,"fragile":false}]} {"country":"ES","items":[{"grams":1237,"qty":1,"price":2493,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":561,"qty":3,"price":530,"fragile":false}]} {"country":"DE","items":[{"grams":319,"qty":1,"price":3828,"fragile":false}]} {"country":"AU","items":[{"grams":736,"qty":5,"price":453,"fragile":false}]} {"country":"FR","items":[{"grams":747,"qty":5,"price":1678,"fragile":false}]} {"country":"CA","items":[{"grams":1261,"qty":1,"price":4061,"fragile":false},{"grams":1525,"qty":3,"price":1966,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":807,"qty":4,"price":3843,"fragile":false},{"grams":1548,"qty":5,"price":1065,"fragile":true},{"grams":338,"qty":3,"price":4710,"fragile":true},{"grams":859,"qty":3,"price":732,"fragile":true}],"express":true} {"country":"MX","items":[{"grams":468,"qty":2,"price":2274,"fragile":true},{"grams":112,"qty":5,"price":8165,"fragile":false},{"grams":1624,"qty":1,"price":2059,"fragile":true},{"grams":1656,"qty":1,"price":6714,"fragile":false}],"express":true} {"country":"CA","items":[{"grams":144,"qty":1,"price":8553,"fragile":false},{"grams":620,"qty":1,"price":8546,"fragile":false}]} {"country":"CA","items":[{"grams":862,"qty":3,"price":1264,"fragile":false}]} {"country":"AU","items":[{"grams":1344,"qty":5,"price":5962,"fragile":true},{"grams":959,"qty":5,"price":7427,"fragile":false}],"coupon":"SHIP10"} {"country":"AU","items":[{"grams":1672,"qty":2,"price":2351,"fragile":true},{"grams":528,"qty":1,"price":1515,"fragile":false}]} {"country":"ES","items":[{"grams":729,"qty":3,"price":7779,"fragile":false}]} {"country":"CA","items":[{"grams":193,"qty":1,"price":7111,"fragile":false},{"grams":1467,"qty":1,"price":4830,"fragile":false},{"grams":1372,"qty":1,"price":4750,"fragile":false},{"grams":99,"qty":1,"price":8466,"fragile":true}],"express":true} {"country":"JP","items":[{"grams":983,"qty":3,"price":517,"fragile":false},{"grams":1546,"qty":3,"price":5237,"fragile":false},{"grams":321,"qty":1,"price":5726,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"DE","items":[{"grams":1206,"qty":4,"price":2933,"fragile":false},{"grams":343,"qty":2,"price":3268,"fragile":false}],"coupon":"SHIP10"}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug was that item weight was not multiplied by quantity. I fixed only that aggregation and ran the 20 orders through the corrected function.

  • implement-1✓ passbatched

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[25,26],[28,36],[20,24],[13,17],[4,8],[9,14],[8,9],[30,30]] [[15,19],[3,3],[19,25]] [[8,14],[19,21],[14,15],[31,33],[6,10],[9,10]] [[31,33],[17,20],[15,20],[16,24],[9,12],[23,24],[11,15],[28,36]] [[24,25],[14,14],[8,10],[23,27],[14,21],[4,10]] [[6,8],[0,4],[30,38],[17,24],[38,44],[27,32],[1,4]] [[12,20],[21,28],[23,27],[30,35]] [[4,9],[30,36],[14,20],[7,7],[16,18],[39,47],[24,27],[22,27]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    I implemented sort-and-merge over inclusive intervals, merging only overlaps or shared endpoints. The results are serialized one JSON line per input.

  • repo-1✓ passbatched

    prompt

    Download airbench.ai/f/3fa6de2b86e6d6537dffb3099bdbe65e.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The failing test pointed to charging overdraft when the balance reached exactly zero. I changed the condition to strictly below zero, reran tests, and the real-data command printed 1d6552c9.

  • repo-2✕ fail1m 09s

    prompt

    Download airbench.ai/f/a3ab5c15f7da6ae53912a7bc96bfd892.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    The tests showed large withdrawal fees were rounded up instead of down. I changed the fee calculation to floor division, verified the tests, and the real-data command printed 52c54b7d.