Building a puzzle game that can grade its own levels
Project Heat is a systemic puzzle prototype about routing warmth through hostile terrain to wake dormant ecosystems. Built solo in Unity over three months, with an editor-side solver that verifies every level is solvable, computes its true par, and reports whether it can be beaten a way I did not intend.
The prototype exists to answer one question
The design document opens with a single question, and everything built since has been in service of answering it: is thermal propagation through environmental constraints inherently satisfying to experiment with and optimise? Not "is this a good game" but "is this mechanic worth building a game around". Framing it that way meant the prototype had a pass or fail condition from day one, and it kept scope honest when features started suggesting themselves.
Heat radiates outward from a fixed source. The player places relay nodes to carry it further and prism nodes to redirect it in a focused beam. Materials in the environment attenuate or block the spread. The level is won when warmth reaches the ecosystem core, and the world visibly thaws as it travels rather than in a cutscene at the end.
I designed and built all of it: systems, tooling, level content, tests and visual direction.
- Heat propagation model and tier attenuation rules
- Material system: stone, glass, metal, ice, frozen river
- Node types: source, relay, prism, ecosystem core
- Level authoring format and loader, then a region and plot campaign structure
- Level Solver: an editor tool for solvability, par and solution space
- Star rating and soft par progression
- Visual direction and art bible
- Automated test coverage for propagation and solver correctness
This is a systems and tooling prototype, not a finished game. Outside the visual direction studies shown here, the in-game art is placeholder and deliberately so: the question above is answered by how the mechanic behaves, not by how it looks, and every hour spent on art before that answer arrives is an hour at risk. What follows is design reasoning and the tooling built to validate it.
A puzzle level is only good if it forces the idea it was built to teach. At ten levels I can check that by hand. At a hundred I cannot.
Propagation that reads without numbers
Heat leaves the source at full strength and weakens as it travels and as it passes through material. The player's job is to find a route. Everything in the design serves the readability of that route, because a puzzle the player cannot read is a puzzle they solve by guessing.
Goals
- Make propagation legible visually, so players reason about space rather than about values.
- Give every material a routing consequence a player can predict before committing.
- Show the world healing during play, not as a reward screen after it.
- Keep systems modular and renderer-independent, so the planned move to isometric 3D does not mean a rewrite.
Four discrete tiers, not a percentage falloff
The obvious model is a continuous falloff: heat starts at 100 and drops by a value per unit of distance and per material crossed. It is easy to write, easy to tune, and it is the wrong choice here.
A continuous value invites the player to optimise the value. It turns a spatial problem into an arithmetic one, and the players who do best are the ones willing to keep a spreadsheet. That is a legitimate design target for some games. It is the opposite of what this prototype is testing.
So heat carries one of four discrete tiers: Strong, Medium, Weak, Dead. Each material crossing costs a whole tier or blocks entirely. The player cannot compute an optimum, only read a state and judge a distance. The cost is real, tuning is coarser and some fine-grained difficulty curves are simply unavailable. The gain is that every routing decision stays a spatial one.
Percentage-based attenuation with numeric readout. Rejected because it rewards spreadsheet play over spatial reasoning, and because a number on screen becomes the thing players optimise regardless of what the designer intended them to look at.
Materials as routing problems
Each material exists to create a specific kind of decision, not to add variety. If two materials would produce the same choice, only one of them ships.
The wall. Forces a route around, and is the basic unit of level geometry.
Reads as an obstacle, behaves as open ground. Teaches the player to check rather than assume.
A relay touching metal inside the heat radius activates through it. Turns a blocked route into a bypass, and is the first mechanic that rewards looking at what a wall is made of.
State that changes during play. The route the player needs may not exist yet, which is what makes restoration a mechanic rather than a visual effect.
The Level Solver
Hand-authoring levels works at ten. The target is a hundred to two hundred, for five to ten hours of play. At that scale the bottleneck is not building levels, it is calibrating them, and there were three specific things I could not do by hand at volume.
I could not be certain a level was solvable at all, only that I had solved it. I could not know the true minimum node count, only the count I happened to find, which meant every par value in the game was a guess. And I could not tell whether a level forced the mechanic it was built to teach, or whether it had a shortcut I had never noticed.
The last one is the real problem. A level that teaches metal conduction, but which can also be beaten by a relay chain around the outside, is not teaching anything. It just looks like it is.
Goals
- Verify that a level is solvable at all.
- Compute the true minimum node count, replacing the par value I had been guessing.
- Measure the solution space: how many ways the level can be solved up to par plus two.
- Identify functional clusters, meaning distinct routing strategies rather than distinct positions.
- Flag unintended solutions that bypass the intended mechanic.
- Live in the editor without touching the shipped build.
Counting strategies, not positions
The solver enumerates every valid node placement and searches depth-first for combinations that reach the core, using a monotonic ordering and frontier pruning to keep the search tractable as candidate counts grow.
Raw enumeration is not useful on its own. A level with two hundred valid solutions sounds rich, but if a hundred and ninety of them are the same route with the relay nudged half a tile, the level has one solution and a lot of tolerance. Those are very different design facts and they need different responses.
So solutions are grouped by obstacle signature: which materials the route actually interacts with, and how. Two placements that negotiate the same wall in the same way collapse into one cluster. What the tool reports is the number of genuinely distinct approaches, which is the number a designer can act on.
One way through. Either a deliberate teaching level or an unintentionally rigid one.
Fine early, a problem in the middle of a campaign.
Wide open. Player expression, or a level that is not asking anything.
Check whether the intended mechanic appears in every cluster, or only one.
What it gives back
Running the solver on a level returns a solvable or unsolvable verdict, the true minimum node count, the size of the solution space up to par plus two, and the cluster breakdown with a plain description of each routing strategy. A Fix Par action writes the computed value straight into the level asset, which removed guessed par from the project entirely.
Figures marked TBC are pending a capture run across the current level set.
Prism solving, liquid water simulation, runtime procedural generation, and in-scene visualisation of solutions were all excluded on purpose and recorded as non-goals before implementation started. The solver covers relay routing only. Scoping it that way is why it shipped in a week and got used, rather than becoming a second project that competed with the game.
Working solo at pace
Three months, one person, 195 commits of tested systems. That output came from a pipeline rather than from long hours, and the pipeline is a design artefact in its own right.
Spec, then plan, then build
Nine dated specification and implementation-plan pairs sit in the repository, from 20/05/2026 to 07/08/2026. Nothing gets built until a spec is approved, and no spec is approved while it still has open questions.
Every spec carries the same four sections: problem statement, goals, non-goals, and open questions resolved. The non-goals table is the one that does the work. It is where a feature gets cut on purpose, with a reason recorded, instead of quietly expanding scope three weeks later.
Tests, because propagation cannot be eyeballed
Twenty edit-mode test suites for a solo prototype needs a justification, and it has one. Whether a heat tier is correct after three material crossings, or whether the solver found the true minimum rather than a local one, cannot be verified by looking at the screen. Both look plausible when wrong. Tests exist exactly where visual inspection fails and nowhere else.
Agent-assisted, with review built in
The build ran through Claude Code with a configured set of reviewer, debugger and test-engineer agents, so code review happens on a team of one. The leverage is real and worth being precise about: it compressed implementation time, and it did not decide the tier model, scope the solver, or cut the features in the non-goals tables. Those are the design decisions, and they are the reason the output holds together.
Deleting the level model
On 07/08/2026 I deleted the flat level data model and its loader and rebuilt the project around region and plot data, sized for a ten-region campaign rather than a list of standalone levels. It cost a week and invalidated working code.
It was the right call for the same reason the solver exists. The flat model worked at eleven levels and would have quietly failed at a hundred, and the cheapest moment to find that out is before the hundred exist.
Writing the look before drawing it
Before any art existed I wrote the visual language as a document, under the working name Amber Hush. Two palettes in constant dialogue: desaturated slate-blues for dormancy, an amber-to-fire spectrum for return. Luminosity carries the entire visual hierarchy, so what is warm is important and what is cold is context. Glow diminishes in soft organic rings rather than hard edges.
The point of writing it first is that it makes the look briefable. A specification precise enough that someone else could execute it is worth more to a project than a mood board I would have to sit next to and explain, and it is the difference between having a taste and being able to direct one.
What it proves, and where it stopped
Eleven levels exist against a target of over a hundred. Onboarding has not been through blind playtesting, so the claim that propagation reads without numbers is supported by design reasoning and my own play, not yet by watching a stranger fail. Solver v1 handles relay routing only: prism solving and liquid water are still manual. The prototype answers its own question in the affirmative for me, which is the weakest possible evidence, and the next milestone is putting it in front of people who have never seen it.
Build the checker before the content
Authoring a hundred levels against unverified difficulty is a hundred levels of rework waiting to happen. The solver cost a week and made every level after it cheaper to trust.
Take the numbers away
Discrete tiers instead of percentages moved the player from optimising a value to reading a space. Showing a number is a design decision, and it decides what players pay attention to.
Non-goals are the scoping tool
Every spec that shipped on time had a non-goals table. Features were cut on purpose with a reason recorded, which is what stopped the solver becoming a second project.