Five rooms before five hundred
A credible building pilot starts small, defines its measurement protocol before intervention, and preserves a real control. This article describes an evaluation pattern, not a completed Smart City Labs customer pilot or measured savings result.

Every operations leader has sat through the same pitch: a platform demo, a sweeping deployment plan, a payback model built on assumptions nobody in the room can check. The pilot that follows covers half the estate, runs for a year, and ends in a report that says "directionally positive." Nothing scales from that, because nothing in it was ever measured tightly enough to survive an argument.
The proposed pattern is small enough to inspect, long enough to cover relevant weather and occupancy variation, and honest enough that a bad result would show up. A real Smart City Labs pilot would pre-register room matching, metering coverage, intervention rules, exclusions, and the analysis window before any operator ran.
Matched rooms, not modeled savings
The standard failure mode of an energy pilot is the baseline. If your "before" number comes from a spreadsheet model, an annualized utility bill, or last year's weather, your "after" number inherits all of that noise. A skeptic can dismiss the result in one sentence, and skeptics should.
A room-by-room design can improve the experiment, but matched rooms rarely differ in exactly one respect. Floor, orientation, equipment, occupancy, guest behavior, weather, and sensor coverage can all confound the result. A defensible estimate needs predefined matching, complete denominators, quality checks, and uncertainty, not meters alone.
Small enough to watch, short enough to trust
Small is not a compromise; it is the instrument. Five rooms can be metered and reviewed by one person more readily than five hundred. At small scale, each supported proposal, configured change, and reported outcome should be inspected individually. Record coverage and recovery behavior must be verified for the specific controller path.
Short matters for the same reason. A pilot that runs a quarter keeps its instrumentation honest and its sponsors engaged. A pilot that runs a year outlives the attention of the people who designed it, and its conclusion arrives after the organization has stopped caring about the question. Set the duration by the physics — enough occupancy cycles, enough weather variation to be representative — and not a day longer.
Scale the routine, not the promise
Here is the part most pilot plans get backwards. The purpose of the pilot is not to prove the technology in miniature so you can buy the big version. It is to build, at small scale, the exact operating routine you will run at large scale — and then multiply the routine.
By the end of a well-designed small pilot, you should know the sensing coverage, policy behavior, escalation path, data gaps, and observed outcomes well enough to decide whether expansion is warranted. Scaling remains a new risk review because equipment, networks, occupancy patterns, and failure modes change with scope.
Contrast that with the estate-wide pilot. It defers every hard question — instrumentation quality, override behavior, failure handling — to a scale where each one is expensive, and then asks you to sign for the full deployment before any of them is answered.
What to ask for
If a vendor, ours included, proposes a pilot, hold it to four tests. Is the baseline a live comparison rather than a model? Is the scope small enough that covered actions and unknowns are individually reviewable? Is the duration set by physics rather than the sales cycle? And does the plan describe the routine you would evaluate at larger scale, not just the result you hope to announce?
Start small. Expand only when the evidence, safety review, and operational routine support it.
Keep reading: A hotel, room by room · The discipline of a low ceiling