Sully
Blog
Opinion

What accuracy actually means on a site walk

Not whether a model labels a crack correctly. Whether the thing you noticed at 10:40am is still in the report on Thursday. Six things that routinely do not survive the trip back to the office.

Jimmy Horan · Founder, Hey Sully · 6 min read
A ceiling corner with a water stain spreading from the wall junction.

Ask most people how accurate an AI tool is and they expect a percentage back.

For site work that is the wrong question, and answering it with a number would be close to meaningless anyway. The failure that costs you is almost never the model misreading a photo. It is that something true, which you observed and understood at 10:40 in the morning, is simply not in the document you send on Thursday.

So the definition worth using is narrower and harder: does the thing you noticed on site survive the trip back to the office?

Measured that way, a folder of photos performs badly, and it performs badly in predictable ways.

Where the loss happens

There is a specific moment when the record degrades, and it is not on site. It is at the desk, days later, when you sit down with a few hundred photos and try to rebuild what you saw.

The photos are all still there. What is gone is the reason each one was taken. You are now grading your own evidence without the context that made it meaningful, at speed, under a deadline. In this segment the pain is well known: a one-to-four hour visit routinely turns into a three-to-five day report. (Directional, vendor-sourced, not our measurement.)

Everything below is a category of thing that reliably dies in that gap. These are illustrative, drawn from the kinds of observations that come up on a walk. They are not customer case studies, and we are not going to dress them up as results we have not measured yet.

1. The pattern you only said out loud

You photograph a cracked render panel. Two rooms later you photograph another. At the third you say, without taking a different photo:

That is the third one of these on this elevation, they are all above the window heads.

That sentence is the finding. The three photos are just evidence for it. Back at the desk, three separate images of cracking, taken eleven minutes apart, do not reassemble themselves into a systemic observation about an elevation. You either remember, or the report says there is cracking in three rooms and misses what it means.

2. The number you read off an instrument

A photo of a moisture meter is one of the least useful images in the folder. Bad angle, glare on the display, and no indication of exactly which part of the wall it was touching. Spoken, it is unambiguous and permanent:

Twenty-four percent at the skirting, twelve a metre up, so it is coming from the floor not the ceiling.

Note that the second half of that sentence is a diagnosis. It is the thing you are actually paid for, it took two seconds to say, and there is no photograph that contains it.

3. The thing that was not there

Absence has no image. A missing handrail, a damper that was never fitted, a smoke seal absent from a riser door, a fire door propped and unlatched with no closer. Nothing in a photo library says this should have been here. You can only catch it by saying it, at the moment you notice it.

This is the category people underestimate most, and it is often the one with the highest consequence.

4. The correction you made thirty seconds later

Real walks are not clean. You say a thing, you look closer, and you revise:

Staining to the ceiling here, looks like the roof. Actually no, there is a bathroom directly above. That is the shower.

Photos preserve neither the first read nor the correction. Written notes usually preserve only whichever one you were writing when you got interrupted. The revision is the accurate version, and it exists for about four seconds unless something is recording.

5. The instruction to your future self

Half of what an experienced person says on a walk is addressed to whoever writes the report, which is usually them:

Flag this for the builder.

If the render is drummy here it will be drummy on the return, check that next visit.

None of that is a defect observation, so no checklist has a field for it. It is workflow, and it is the difference between sixty photographed items and the eight that generate action.

6. The thing in the background

You film a ceiling to record water staining. In the corner of the same frame, for two seconds, is a distribution board with its cover off.

You were not looking at it. You would not have photographed it. But it is there, in the footage, permanently, and it is findable later when someone asks whether anything was noted about the switchboard. Continuous capture records things you were not paying attention to. Photographs, by definition, only record what you framed.

The pre and post case

The clearest version of all of this is the corridor pre-and-post pair. You baseline every adjoining property before the works, then walk the same properties again afterwards and produce a difference.

The value of the baseline is entirely determined by whether it recorded what someone will dispute two years from now, at which point nobody can go back and re-take the photograph they did not think to take. A narrated walkthrough is a much better baseline than a photo set for exactly the reason above: it captured things you were not paying attention to, and it captured why you thought what you thought.

This segment is also the one where video is welcomed rather than resisted, because a moving, continuous, time-stamped record is stronger evidence than a folder of stills.

Why we do not publish an accuracy percentage

Because for this problem it would be theatre.

A single number would average over wildly different question types, it would come from our own test set, and it would tell you nothing about whether the specific thing you noticed on your specific job made it through. Anyone quoting you a headline accuracy figure for open-ended site work is selling something.

What we do instead is make the draft checkable. Every line Sully writes points back to the frame it came from and the words that were said at that moment. You are not asked to trust a summary. You are asked to glance at the citation, confirm it in a second, and move on. When it is wrong, you can see that it is wrong immediately, and where it went wrong.

That is a weaker claim than "99% accurate" and a much more useful one. It also happens to be the posture that professional bodies are converging on for AI in signed work: a named human reviews, the review is on record, and the trace exists if anyone asks.

The draft is never sign-ready. You review it, you correct it, you sign it.

The practical version

You do not need us to get most of this benefit. Talk while you capture. Say the reading, say the cause, say what is missing, say what needs to happen. The habits are written up in how to narrate a site walk, and they work with any phone and a voice memo.

What we add is that we take the walkthrough and the narration as a single thing rather than two files, and give you back something structured and searchable, where every line traces to its source. Not a better camera. A record that still contains what you knew on the day.