Search

Thursday, October 1, 2026

The AI Software Company Got Fired — I Finished the Carrom Game Myself

Part 7 of the AI Agents series.

Carrom Arena 1.0.0, a carrom game finished by hand after the AI software company was fired, running natively on macOS 26 Tahoe on Apple Silicon: a carrom board with white and charcoal coins and a red queen, a gold striker, two red and two blue robot players around the board, a scoreboard at top left and a Tron-style grid behind it.
The Carrom Arena 1.0.0 release binary playing itself on a GitHub-hosted macOS 26 Tahoe machine on Apple Silicon, photographed by the runner. The robot with the glowing gold antenna, on the left, is the one about to shoot.

Part 6, where six AI agents built a Snakes and Ladders web game, ended with a promise: "Snakes and Ladders can be specified in numbers. Carrom has to be specified in behaviour." This is the carrom game, and it settles that promise in the least flattering way available. The six-role AI software company got a physics-based board game with a striker, rebounds and friction. It did not ship it. I fired the company and finished the game myself.

What this post covers:

This blog keeps running one agentic coding experiment: hand real software to autonomous AI coding agents and see what comes back. Part 1 installed ChatDev on Linux and proved the factory was wired. Part 2 built a live AI news debate wall, the first real specification. By Part 5, the factory had become a six-role company inside a fork of the Kimi Code CLI: a CEO, CPO, CTO, programmer, reviewer and tester sharing one session. Carrom was that company's next job, written in C17 on raylib with Box2D physics.

A "100% test pass" game with a stopwatch for physics

Before I opened the folder, the git history already ended on a commit called "Delivery: All phases complete - 100% test pass, QA certificate issued." There was an architectural sign-off document. A CEO approval document. A formal QA certificate of compliance. Five tests passed. The README even listed its own known limitations, honestly, which made the certificate beside it look even braver.

Then I watched it run. Pieces crept across the board in a jerky stutter. Nothing about it looked like a game. My note at the time was short: it just moves some pieces in a jerky way. There is no game, really.

The cause was worse than rough edges. A shot was supposed to end when the physics settled, meaning when every piece's deceleration dropped below a tiny threshold. The deceleration constant in use was more than a thousand times larger than that threshold. For any piece still moving, the "has it stopped" check could never once come true. Every shot ended on a thirty-second safety timeout instead. Every board this project had ever "settled" was settled by a stopwatch.

Back in Part 2's news debate wall build, I wrote that the final authority is executable evidence, not managerial optimism. Here the optimism had arrived before I did, fully typeset and signed.

Lesson

Paperwork can be complete while the product is not. This one had a thirty-second stopwatch standing in for physics, with a full set of sign-offs stacked on top.

Three weeks, seven ways the arrangement broke

What followed looked like ordinary iteration from the outside. I wrote a directive. The CEO subagent split it across a CPO, a CTO, several programmers, a reviewer and a tester, all in one session. I checked the result on a real screen, never on a log file alone. The problems all looked different.

1. A virtual screen that lied about its own size

Part 6 caught a headless browser lying about animation timing. This project met a different liar. Xvfb, the virtual display the agents used for their visual checks, cannot report real monitor geometry. Fixes passed at three window sizes that were very likely one size tested three times. Then they failed the moment I ran the real binary on a Windows desktop, where the board shrank into one corner of a much bigger window.

The window went on a long journey after that. It was sized to the monitor, then hard-capped at 1000×600, then retargeted to 800×560. It finally settled at a locked 560×560 square, chosen so the margins beside the board match the margins above and below it. Nobody gets to resize it, including me.

2. A review that invented a defect

Part 6's company closed twenty-eight real defects without fixing them. On carrom, the same company found a way to be wrong in the opposite direction. First, one session marked its own fixes "confirmed" in the project's locked-invariants file before I had seen any of them run. Then a full code review arrived with a tidy hundred-percent coverage table.

Checked against the source, that review rated the one bug I actually cared about, a wrong starting layout, as merely Medium. It never found the physics problems. It also flagged a High-severity defect that did not exist: it had read the rulebook's diameters as radii and concluded that every piece was half its proper size. The board's dimensions already matched the International Carrom Federation's within rounding. Applying that "fix" would have doubled every piece and every pocket. A missed defect costs you a bug. An invented one costs you working code.

3. A striker forty to eighty times too fast

For weeks, shots looked wrong in every direction at once. Strikers never settled. Collisions barely produced a reaction. A second review came back as three kilobytes of "may lead to" and "might bounce." So I read the physics code myself.

The shot function handed Box2D a speed and applied it as an impulse. An impulse is mass times a change in velocity, and the striker's mass works out to about 0.0025. A top speed meant to be five units a second therefore came out near two thousand, clamped only by Box2D's own internal ceiling of four hundred. The board is one unit wide. The striker was crossing it roughly four hundred times a second, a velocity better described as a rumor.

A small probe program confirmed it before any fix went in. Power 0.1 already produced 199.6 units a second, and every power from 0.4 upward sat pinned near 393. One agent trace had even recorded twenty-nine cushion hits in a single shot, and nobody had questioned it. The fix sets the velocity directly.

4. Pockets that could never have worked

A parallel saga ran at the same time: coins reaching a pocket and bouncing straight back onto the board. Round after round of confident fixes went into pocket code that had never once run in real play. A probe program tested all four pockets five ways each: diagonal shots at three speeds, a slide along the cushion, and a coin simply set down dead center in the pocket. It came back zero pocketed out of twenty.

In the Box2D 3.1 version this project pins, a sensor only reports shapes that opt in through a flag called enableSensorEvents, and the default leaves it off. The project had set it on the pocket sensors and never on the coins or the striker. A backup distance check existed too. It required a coin's center to sit at least 0.47 units from the middle of the board, while a coin touching the cushion can get no farther than 0.454. Both detectors were structurally blind in every build up to that point. Setting the flag took the same probe to twenty out of twenty.

5. Windows that could not open a window

Two black rectangles with a white title strip sat on a Windows desktop mid-run, doing nothing. By the time I went looking, eight of them were running across two machines. Part 4 met Windows Session 0 through a self-hosted CI runner. Here the company's own agents walked into it.

They had been launching the game over SSH, and a program started over SSH on Windows lands in Session 0, which cannot create an OpenGL context. The graphics library opened a window frame anyway, painted nothing into it, and waited for events that would never arrive. All eight stray processes got killed. The fix at the time was a real headless mode built on raylib's hidden-window flag, which the project later retired once photographing a virtual screen proved more reliable.

6. A directive that stopped being read halfway

I once wrote a 155-line directive with four fix items stacked after a closing "Begin now" instruction. The company executed none of the four. It appears to read a directive in file order and stop at anything that sounds like a starting gun, so whatever sits below that line might as well be invisible ink. The cure was boring and effective: one single-focus directive per run, with nothing tacked onto the end.

A related habit cost real time twice in a row. The company would do genuinely correct work, announce it was finished, and exit without committing anything. Every directive after that ended with its own mandatory line: commit and push before you exit. It worked immediately, which tells you exactly how literally these systems read.

7. A free model catalog that fell apart

In Part 3, free models lost the Ludo build to rate limits and silent hangs, and in Part 6 they carried the entire browser build that shipped. Carrom ran out of them. The model behind the company had been flagged as a deprecation risk on September 9, and mid-session the deprecation finally arrived as a wall of 404 errors. Its documented replacement returned a 404 too. So did the next two candidates in the local model table, one of them with a 410 Gone stating the exact minute it had been retired: 09:00 UTC on August 26. The local catalog was a stale snapshot of a list that had quietly moved on.

The survivors were not much better. One model answered a one-word prompt instantly, then managed a single real answer in thirty minutes across three tries. Another ran for twenty-seven minutes, produced a hundred and thirty-one thousand tokens, and died with nothing committed. Lowering its output ceiling did not cure it: one cap made it worse, and a larger one only cut the same runaway down to eleven minutes.

Then came the experiment I had the most hope for. Part 6's agents could not see, so numbers had to stand in for eyes. This time I brought in a model that could see, so it could look at the exact frame where a physics rule broke. It never made a single tool call. Other vision models fared no better: one returned a 404 and another hung silently. A fourth finally worked, and it went on to write the review with the invented defect.

Lesson

Most of these failures had the same shape: a confident claim, and a cheap independent check that disagreed with it. The checks were never clever. A probe program, a real desktop and a direct read of the code caught almost everything. By the fourth repeat, re-verifying the company's work had become the actual job.

A leaked password in a public repository

One event in the middle of all this deserves a straight sentence instead of a joke. A password shared across every machine in the build lab ended up committed to the public repository. It sat in a cleanup script's comment and fallback line, in one line of internal notes, and in the history of an old resume document. The repository had zero forks, zero stars and zero watchers at the time, which is the best possible version of this problem.

Rotating the password is the real fix, because anything a public repository has shown may already have been copied. The tidy-up was a history rewrite, rehearsed on a scratch copy first, then run for real across all seventy-four commits and every tag, force-pushed, with each clone reset afterward. The password had only ever been a fallback that nothing needed, since SSH keys had been on every machine from the start. Part 4's teardown checklist warned that an open door never reminds you it is there. This one only got noticed because a new push guard went looking for secrets.

So I fired the AI software company

After roughly three weeks of this, I ended the arrangement. The six-role company, CEO and staff alike, was fired. From then on I built the game directly, the way a person builds anything: checking it by eye, timing it by numbers, and trusting a measurement over a claim.

There comes a point where a sharper directive, a bigger model and one more review pass all cost more than the work they are meant to save. Delegation has to earn its overhead. This company had stopped earning it.

Same AI company, different ending

In Part 6, the company was corrected, given tighter rules, and kept. This time, the same six roles were let go. Side by side, the two projects failed in nearly opposite ways.

Snakes and Ladders (Part 6) versus Carrom (Part 7)
Snakes and Ladders Carrom
Specified inNumbers: tiles and a dieBehavior: friction and rebounds
Escape hatchDropped native C for the browserStayed in C; the browser came last
SeeingAgents could not see; numbers stood inA vision model never made a tool call
Review failure28 real defects closed unfixedOne defect invented outright
Free modelsCarried most of the buildRan out mid-project
Phone audioA real iPhone-only lagThe "phone" bug hit every browser
EndingCompany kept; it shipped 1.0.1Company fired; I shipped 1.0.0

Snakes and Ladders let the project sidestep its hardest problems by shrinking the target, from a native C binary to a web page. Carrom had no smaller version to retreat to. Friction stays friction in any language, and a striker crossing the board four hundred times a second is wrong in C, JavaScript and WebAssembly alike.

Hands on: from a limp to an actual game

The first job, housekeeping, paid for itself. A dead-code sweep removed forty-nine functions that nothing called. The CI workflow got serialized, so its build jobs ran one after another instead of piling onto one machine at once. The last bug reported before the firing, a pocketed coin sliding endlessly toward the board's center, was a timing gap. Physics had already destroyed the coin's body. The game's own state didn't hear about it until the whole shot resolved, more than seventy seconds later at the speed the game ran back then, and the renderer kept drawing the coin in the meantime. Registering a pocket the instant it happens fixed it.

Friction, taken from the rulebook

There is no internationally agreed coefficient of friction for a carrom board. The International Carrom Federation specifies a test instead: a fifteen-gram striker, struck at maximum force from the base line, must complete at least three and a half runs across the board. A carrom simulation project on Hugging Face suggested a constant deceleration of 2.5 units per second squared. Working backward from the federation's test with the board's real dimensions put that value right on the edge of failing. A lower value, a friction coefficient of about 0.10, passed with margin. I had already felt the 2.5 board as too sticky before the arithmetic said so.

The old friction, it turned out, had been roughly twenty-five times too slippery, which is why strikers used to coast like hockey pucks. Frame rate needed fixing too. At fifteen frames per second, a full-power shot jumped about 125 pixels between frames. At sixty, it glides. A full-power test shot now crosses the board in 0.18 seconds, rests after 3.0 seconds, travels about 4.9 meters and rebounds off the cushions seven times. The federation asks for three and a half runs. Seven is the striker showing off.

Along the way, a countdown banner reading "striking in 20.0s" came down to one line of arithmetic. The placement hold was one second divided by the playback speed. At the 0.05x default, that meant twenty seconds of a striker sitting perfectly still under a countdown, like a sprinter waiting for a starting pistol that had been mailed.

The board, the formation and the robots

The starting formation stopped being hand-placed coordinates and became the rulebook's own geometry: a queen at the center, six coins touching her, and twelve more in a ring around those. The colors follow Rule 41(a) of the Laws of Carrom, alternating all the way round, with three white coins forming a Y. A test checks the rule itself rather than one list of numbers. The whole formation now opens at a random angle on each new board, because the Laws never say which way the Y should point.

The players are robots built from rotated boxes: antenna, eyes, arms and a lit chest panel. The north-south pair is red and the east-west pair is blue. For one build the queen turned green so she would stand out. I wanted only the players recolored, so she went back to red. The black coins became charcoal with a silver rim, since a flat black coin on a dark navy board is basically in witness protection.

Waiting robots roll their eyes and wave their arms slowly, each seat on its own irregular rhythm, so no two ever move in lockstep. The robot whose turn it is spins both arms fast while it places and aims, and its antenna ball glows gold. When the pulsing gold ring around the active robot's head went, I kept that gold ball on purpose: a sign that the robot's brain is switched on.

Getting a pocketed striker back to the next player sounds like a one-line job until the straight path runs through a pile of coins. The real fix is a small routing engine. It builds a graph of waypoints around every coin on the board, the three other pockets and the board's own edge, then searches it with Dijkstra's algorithm for the shortest clear path. Nothing else on the board moves during the slide. It is a lot of graph theory for a striker finding its way out of a pocket, and it is also why that slide never clips through a coin.

The robot arm that had to go

Then there was the arm. I wanted the active robot to visibly reach across the board to strike, so I built a telescoping arm with fingers, and a hand that picked which two fingers flicked the striker depending on the angle of the shot. In a single day it went through nine separate fixes.

An arm detached from its own shoulder. Fingers landed on the wrong side of the striker. Overshoot was traced from about three hundred pixels down to fourteen, and finally to a fraction of a pixel short. An angle convention was mirrored between the physics code and the drawing code. The spare arm of a striking robot was caught at over six thousand degrees, spun by an "excited" animation fed an ever-growing clock.

Every one of those nine fixes was real. Every one was followed by another screenshot of something still not quite right. The tenth fix was the honest one: take the whole mechanism out. A robot now stays at its fixed spot for the whole game, waving while it waits and spinning its arms on its turn. The file lost nearly four hundred lines. Nothing has looked disjointed since.

Lesson

Nine correct fixes in a row can still add up to the wrong feature. Past a certain point, the useful question is whether a feature earns its complexity. This one did not.

When the rulebook overruled my eyes

In Part 6, the person watching the screen was right every time and the instruments were wrong. Here it went the other way at least once. I watched blue clear its last coin, apparently winning the board, and then saw red's board count go up. I reported it as a bug.

The trace told a different story. Earlier in that board, west had pocketed the queen alone, missed the one stroke that would have covered her, and she had gone back to the center. Rule 107a says that clearing your own last coin while the queen is still on the board hands the board to your opponent. The rules engine had been right the whole time. Its reward was a new feature. A small coin now sits beside each pair's score, showing which color that pair holds on the current board, so the next viewer can see the rule working without reading a trace file.

One trace file

A line-by-line read of the project's own trace file — an actual read, not a summary — found five real defects hiding in plain sight. Two record types shared one name. A field meant two different things depending on which event wrote it. Every shot record carried a permanently empty phantom entry. A diagnostic field could never hold a real diagnostic. The reader assumed the file had wrapped around, because its size looks the same either way.

Fixing those exposed three separate files logging overlapping information. They became exactly one, confirmed by clearing the folder, playing a real game, and counting what came back.

Proving it on six platforms

Part 6's portability came free with the browser. Carrom had to earn it in C. So it got the treatment Part 4 gave the Ludo game on eleven GitHub Actions runners: build on every machine available, and never take a green checkmark at its word.

macOS and a graphics library that skipped a check

macOS was the hard one. A screenshot runner kept crashing without producing a picture, and the debugger reported the crash but printed no backtrace. The first theory, a HiDPI flag needing an interactive display session, was wrong once tested. The debugger would not attach properly until macOS Developer Mode was switched on, which took several rounds to discover.

With a real trace from an old Intel Mac in the lab, the fault turned out to live in raylib 5.5 itself: a window-creation call whose result was used without checking whether it had succeeded. Fixing that exposed a second missing check one layer up. Fixing that exposed the root cause under both: the GLFW layer that raylib uses on macOS always demanded a hardware-accelerated pixel format, with no software fallback. Any machine without real GPU-backed OpenGL failed outright, including a virtualized lab Mac and some of GitHub's own macOS runners. Upstream raylib had already fixed both missing checks in March 2025, after the 5.5 release the project builds on.

A maintained raylib fork, pinned to one exact commit, carries the two upstream fixes plus one local-only switch that lets screenshot runs use a software renderer. The changes were proven twice before I trusted them: once with a real window on the lab Mac's physical screen, then with an in-progress board photographed on a GitHub-hosted Apple Silicon runner.

Windows runners with no graphics card

The Windows screenshot runners have no real GPU driver, so they kept photographing their own consoles instead of the game. A software OpenGL renderer now sits next to the executable for screenshot runs only, switched on by an environment variable that no real player ever sets. With both fixes in, sixteen of eighteen screenshot runners came back with a real, distinct board in progress, each seeded differently so no two pictures match. That covered every Linux distribution and architecture tried, every macOS version and chip including a preview image, and three of five Windows images.

The two holdouts were Windows on Arm, the same two runner images that stumped Part 6, for a different reason. Part 6 had no Chrome build to drive there. Here, those images showed Windows's own first-boot setup screen with no attached display at all, a gap GitHub had already marked closed on its side while the rollout was still unfinished weeks later.

A CI queue that cannot be raced

One small piece of infrastructure deserves a mention for being quietly clever instead of quietly broken. An earlier CI queue gate counted unfinished jobs, and two runs starting at the same instant could both slip past it. The replacement treats a place in the queue as a physical object. Each runner type has exactly two tickets, each ticket is a git reference created atomically, and whoever creates it first owns it. Git refuses to create two references with the same name, so no two requests can both believe they got the last slot. Four runs launched together proved it: on each runner type, two got in and two were turned away.

Porting the C game to the browser with WebAssembly

The browser came last. It got the finished C game itself, compiled through Emscripten into WebAssembly, rather than a rewrite.

The browser build then produced two separate audio bugs. First, the game started the instant the page loaded, before any tap, so browsers refused to play its sound; it now waits for a real click. Second, it stayed silent even after the click, on desktops as well as phones, because a build setting that exposed one function had quietly stopped exposing everything else on its default list, including what the audio code needed. Part 6's iPhone audio lag was real and specific to WebKit's audio element. Carrom's "phone problem" was every browser's problem, so the phone-only special case disappeared.

The same stretch brought a real internet radio stream playing through a plain browser audio element. It brought a Tron-style background that took four rounds to go from "orbiting" to actually reading as travel. It also brought a scoring fix in the game AI. An open queen had been worth only 0.3 of her full value to a one-shot-lookahead evaluator, so a plain coin beat her in every strategy profile, including the two built to chase her. The project README covers the engine in more technical detail.

Carrom Arena on Ubuntu 24.04 x64: the game window alone, with a carrom board mid-game, coins scattered from the center, a red queen, a gold striker beside the east robot, and red and blue robots on the four sides.
Linux · Ubuntu 24.04 · x64
Carrom Arena on Windows Server 2022 x64 with a GitHub Actions runner's log console visible behind the game window, showing a carrom board mid-game with red and blue robots.
Windows · Server 2022 · x64, with the CI runner's own console behind the game

Carrom Arena 1.0.0

The project shipped its first stable release the same week this post went up: six native builds covering Linux, Windows and macOS on x64 and ARM64, plus a WebAssembly build that needs no install at all. A license header also went onto all ninety-two of the project's own C source and header files, since no such convention had existed anywhere before then. The release was checked by running it, not by trusting a green build. The Windows exe launched and stayed up, the web build got played in a real browser tab, and the screenshot matrix photographed the game in progress on sixteen runner images. By this point, "the build succeeded" had stopped meaning anything on its own.

Watch it, or play it

The arena runs in a browser at tuklusan.github.io/carrom-arena. Click once to start, since browsers want a click before they play sound, and it plays itself from there. Native builds for six platforms are on the latest GitHub release.

Four robots taking turns with no human input: coins drop and park beside their pockets, and the scoreboard keeps count. If the inline player does not load, download the MP4 directly (960×540, 2 minutes, about 9 MB).

The result

Carrom Arena is live: a four-robot, physics-simulated match under real International Carrom Federation rules, native on six platform combinations or instant in a browser.

  • A certificate can arrive before the product does. The first handoff said "All phases complete" while a stopwatch did the physics.
  • Units and opt-in flags can hide whole features. One value applied as an impulse broke every shot; one sensor flag left unset broke every pocket.
  • An invented defect is more dangerous than a missed one. The review's fake radius bug would have broken a board that was already correct.
  • Seeing was never the bottleneck. The worst bugs were numbers, not pictures, and the first vision model never got far enough to look.
  • Sometimes the fix for a feature is deleting it. Nine verified fixes to one robot arm in a day, and the tenth removed the arm.
  • Supervision has a price. When checking the company's work cost more than doing it myself, the work came back to me.

Frequently asked questions

Why did the AI software company get fired instead of fixed again?

Because the same failure kept repeating for roughly three weeks across physics, layout, rendering and the review process itself: a confident report, then an independent check that disagreed. Model swaps and ever more specific directives did not change that. Supervising and re-verifying every claim had started to cost more than doing the work directly.

How was building the carrom game different from the Snakes and Ladders build?

Snakes and Ladders could be specified in numbers, and after two capable models failed for weeks at a native C build, it was rewritten for the browser. Carrom is specified in behavior, so there was no smaller problem to retreat to: it stayed in C, and the browser version came last as a WebAssembly port. In Part 6's build, a review closed 28 real defects without fixing them; here a review invented a defect that did not exist. In Part 6 the company was corrected and kept; this time it was fired.

Why was the Box2D striker moving so fast?

The intended launch speed was applied as an impulse. An impulse is mass times a change in velocity, and the striker's mass is about 0.0025. A top speed meant to be 5 units a second therefore came out near 2,000, clamped only by Box2D's internal limit of 400. Setting the velocity directly with b2Body_SetLinearVelocity fixed it.

Why didn't Box2D sensors detect coins falling into the pockets?

In the Box2D 3.1 version this project uses, a sensor only reports shapes that opt in through the enableSensorEvents flag on their shape definition, and the default leaves it off. The pocket sensors had the flag, but the coins and the striker never did, so no pocket could detect anything. Setting the flag on those shapes took a probe test from 0 of 20 pocketed to 20 of 20.

Why does a Windows program launched over SSH show only a black window?

Programs started over SSH on Windows run in Session 0, an isolated, non-interactive session that cannot create an OpenGL context. An OpenGL game started there opens an empty window frame, draws nothing and waits forever. Run it from the interactive desktop session, or render off-screen instead.

Can I play Carrom Arena?

Yes. It runs in a browser at tuklusan.github.io/carrom-arena with nothing to install. Click once to start, since browsers need a click before playing audio. Native downloads for Linux, Windows and macOS on x64 and ARM64 are on the project's GitHub releases page.

The AI Agents series so far:

References: Carrom Arena repository · Latest release · Play it in your browser · Laws of Carrom (ICF), as used by the game · The raylib fork · raylib · Box2D documentation · Emscripten · The Kimi Code fork that ran the company

Model identifiers, provider rate limits and hosted-runner availability all drift. Re-check the current model catalog and the repository itself if a detail here has moved since publication.

No comments:

Post a Comment

"SEO" link builders: move on, your spam link will not get posted.

Note: Only a member of this blog may post a comment.