SOLREIGN//DEV BLOG
Dev Blog #12

How the Station Was Built With AI

·The person building SOLREIGN

STATION DIRECTIVE /// DEVLOG BRIEFCLEARANCE: PERSONNEL
A reader asked how a one-person Space Station 14 fork uses AI day to day. Here is the workflow, what it is good at, what it is bad at, and how a solo developer starts.

Someone on the Discord asked, plainly, how a one-person Space Station 14 fork actually uses AI day to day. The honest answer is that the workflow is unglamorous, the wins are small, and the things that broke were the things I assumed the model would get right. This post is the long version of that answer, for the player who asked and for any solo developer wondering whether the trade is worth it.

The station you play on was built with AI. It was also built without it. That is the boring truth, and the boring truth is the one that survives contact with a long-running codebase.

The shape of the work

A typical day looks like this. I pick one feature, usually a small one. I write a short spec for it. The spec lives next to the code, in a folder called design docs, and it stays there after the code lands. The lore of the station, the in-world reason a feature exists, lives in a separate folder of canon notes. Specs are how I remember what I was thinking six weeks later, and how a model gets pointed at the right question instead of the wrong one.

Once the spec is on paper, I split the feature into slices. A slice is a change I can describe in one sentence and test in one command. The empty-round autopause guard was one slice. The audit-chain decorator-count pin was one slice. The arrivals beacon was one slice. Each ships on its own branch and writes its own changelog line. Big-bang PRs are how a model quietly changes something the test suite never looked at, and you do not find out for a month.

The slice is what the model works on. I hand it the spec, the slice description, the canon note, and a small number of files. I ask for a diff. I read the diff line by line. I do not accept a patch I cannot explain. The number of times a diff has looked reasonable and turned out to do something the spec did not ask for is non-zero, and the cost of catching it at the diff is five minutes. The cost of catching it in production is the rest of my week.

After the diff lands, the test suite runs. The test suite is the part of this I do not let the model write on its own. The model proposes. I author. The rule is that every test on the production path exists in a form I could have written by hand, and the model cannot talk me out of that rule by sounding confident. Confidence is not a verification.

If the tests pass, the slice goes to a local server. I play it. This is the gate that catches the most bugs. A Space Station 14 round is a stress test that no model can simulate, because the whole point of the round is that four humans do four unexpected things in the same five minutes. If a slice survives a playtest, it survives for a reason. If it does not, the slice is not done. The last step is the changelog. Every shipped slice writes a line that says what shipped and what it broke and what did not. The changelog is the part of this the player actually sees, and the player deserves a sentence that is not a victory lap.

What AI is good at in this codebase

The first is reading. A model can chew through a 4,000-line YAML config and tell me which job spawns where and which one was deprecated in v286. That kind of work used to take me a Saturday. Now it takes a coffee.

The second is refactor drafts. If I tell a model to rewrite a 200-line method so the responsibility for curtime lives in one place, it gives me a diff. The diff is usually 70% right. The remaining 30% is the part where it loses track of a side effect, and that 30% is what I am paid to read.

The third is the boring middle of a feature. Once a spec is on paper and the slices are written, the slice implementation is mostly typing. The model is fast at the typing. I am still the one who decides what to type.

What AI is bad at in this codebase

The first is the part of the work I cannot describe. The lore reasons a feature exists are not in the codebase. They live in the canon notes, and even when the model has the canon notes open, it does not always understand the in-world reason a given mechanic is the way it is. The model is good at the what. It is bad at the why.

The second thing is judgment. A model will happily suggest a fix that passes the test suite and breaks the round. A green test suite is a necessary condition, not a sufficient one. The sufficient condition is the playtest.

The third thing is steady-state discipline. If I let a model loose on a feature without a spec, it produces code that is plausible and wrong in ways I cannot detect by reading. A 2,000-line diff is not a diff. It is a betrayal, and the model did not know it was betraying me, which is the worst part.

How a solo developer starts

If you are a solo developer wondering whether to bring a model into your own project, the first week is the only week that matters. Pick one feature, not the project. Write a one-paragraph spec for it. Split the feature into slices, each a single diff, a single test, a single changelog line. If a slice is too big to describe in one sentence, it is too big to be a slice. Let the model write the diff for one slice. Read every line. Reject any line you cannot explain. Run the test suite, playtest the slice, write the changelog line. Then repeat.

The cadence is the point. The model is a tool. The cadence is the thing that makes the tool load-bearing.

A small note on the year

A lot has been written this year about whether AI is going to replace the people who build games. The honest answer from inside one small station is no. The fork is possible because a person decided how the station should feel, wrote that decision down, and then asked a tool to help with the typing. The tool was useful. The decision was not the tool's.

If you have read this far, you are the reason the fork exists and the reason I write these posts. Thank you for playing. Thank you for asking. The next post will be a thing that broke, because something always breaks, and the changelog line for it is overdue.

Station Boarding Advisory

Ready to step onto SOLREIGN?

SOLREIGN is an independent, persistent Space Station 14 server. Every other station forgets you at round end — this one keeps a file. Free to join, airlock open.

← All Dev Blog entriesHow to Join →