Finding Ships in Radar, Without a Neural Net
Finds ships in free Sentinel-1 radar and measures them. 76% accurate on vessel type over 326 vessels, checked against AIS.
Draw a box over water, pick a Sentinel-1 pass, and get every ship in it with a measured length, beam and heading, plus a cargo / military / other call. Runs entirely on your own machine, no account anywhere. Exports are a GeoJSON of the contacts plus a decibel GeoTIFF, both of which drop straight into QGIS or ATAK.
Type accuracy is 76%, measured against the broadcast dimensions of 326 AIS vessels. It started at 56%. Getting from one to the other is what this is about.
Try the browser preview — no install, no account. It runs the same CFAR, the same measurement and the same classifier in your own browser, and pulls the radar straight from the archive, so where you look never reaches my server. The desktop tool is still the real one: this preview has no AIS correlation, no video export, and no memory between passes.

Why radar instead of a photograph
Optical satellites need daylight and a hole in the clouds. Synthetic aperture radar needs neither, which is the entire reason it exists. A ship is a corner reflector sitting on a surface that scatters almost nothing back, so vessels show up as bright points on a dark field. That contrast is the whole detection problem, and it is why this works without machine learning.
The data source is Microsoft Planetary Computer's sentinel-1-rtc collection. It is anonymous, 10 metre, north-up UTM, and cloud-optimised. The alternatives all fail on access rather than quality: earth-search's GRD bucket is requester-pays, Copernicus Data Space wants an account and enforces quota, and ASF wants an Earthdata login. RTC's real limit is that it is only processed where a terrain model exists, so there is no mid-ocean coverage. That does not matter. Ships are in ports and straits.
Detection is CFAR
Constant false alarm rate detection compares each pixel to the statistics of the water around it rather than to a fixed brightness. A pixel is a candidate when it exceeds its own local background by enough that the odds of speckle producing it are below your tolerance.
The local mean and variance come from integral images, read over a 31 pixel window with a 7 pixel guard band around the pixel under test, and the threshold is that mean plus k standard deviations.
The trap is memory. A naive implementation builds six float64 integral images the size of the input, which is 2.5 GB at 4096 by 4096. Tiling the detection fixed that and raised the practical area cap from 1,500 to 12,000 square kilometres.
The breakwater ate a day
A seawall fires the detector as a dashed line, and every dash measures like a 40 metre trawler. Two global filters were built and thrown away.
Connected components on aspect ratio was flawless on synthetic data and masked nothing real, because speckle bridges the wall into one blob 15 km long and 9 km wide. A rotate-and-open directional line filter found the wall and also every sidelobe streak coming off a bright hull, then masked the hull sitting at the centre of its own starburst. Detections went from 59 to 9, almost all of them real ships.
Both filters passed their tests. The tests were the problem, not the filters. Nothing catches this except rendering the mask and looking at it.
What survives is a per-blob probe that asks whether a single detection sits on a locally linear structure. It leaves about four bright spots on the Long Beach wall at any probe length. That is documented in the interface rather than hidden.
Classification is a database, not a model
Six measured features contribute a log-likelihood against a class table: length, calibrated beam, the length-to-beam ratio, where the superstructure sits along the hull, how far the brightest pixel rises above the local sea, and the cross-polarised return rank within the scene.
The tool reports the split instead of picking a winner. A 155 metre contact is a destroyer or a feeder container ship, and no amount of confidence styling changes the fact that length alone cannot separate a warship from a merchant of the same size.
The validation harness found the real bugs
NOAA Marine Cadastre publishes AIS as free daily archives, no account, roughly 395 MB a day. Scoring the pipeline against it found three defects that staring at the map never would.
- Beam ran about 1.9x too wide, from the radar point spread plus the detector's own morphological closing adding roughly a pixel per side to a three pixel beam. That correction later turned out to be treating a symptom: see below.
- A 3x3 closing was fragmenting large vessels, reporting a 59 metre piece of a 333 metre tanker as a fishing boat.
- Two real classes were missing entirely. Wind farm installation vessels are 210 by 20 metres, far too slender for any cargo class in the table.
Fixing those took type accuracy from 56 to 63 to 69 percent.
The beam correction was measuring the wrong thing
A single multiplicative bias should travel between ports. This one did not: it wanted 3.36 at Skagen and 2.29 at Long Beach, a disagreement of 1.47x that no constant can absorb. The reason is that the old beam was a spread statistic over a point set, so anything the blob contained widened it — sidelobe arms, azimuth ambiguity ghosts, a second hull rafted alongside. On a very bright target that contamination is the majority: a 276 by 46 metre tanker arrives as a 340 pixel blob of which barely a third is ship, and the quantile then measures the flare. That is how 46 metre hulls came back at 158.
A hull is not a point set, it is a solid. At any station along its length the pixels crossing it form one unbroken run straddling the centreline, and flare is either detached from that run or far thinner. Walk the hull, take the run at each station, use the median. The two ports now want 1.74 and 1.64 — a gap of 1.06 — and the correction factor falls from 1.93 to essentially none.
Quoting a spread instead of a verdict was the other half. A binary "trusted / not trusted" flag was built and withdrawn, because the thing that decides whether a beam is good is the true beam, which is exactly what is not known at runtime — it was calling correctly measured slender tankers untrustworthy. Every contact now carries a real plus-or-minus instead. Combined with the class prior, median beam error went from 17.4 m to 6.6 m, and the share within 10 m from 29% to 63%.
The harness had bugs too
It fed the classifier an already-corrected beam figure, double-corrected it, and reported 53% when the truth was 69%. Its weight sweep then recommended disabling three features for a one-vessel gain on a sample of 45, where the binomial noise floor is about three. It now computes that floor and refuses to recommend anything inside it.
A broken measuring stick is worse than no measuring stick, because you believe it.
What it cannot do, stated plainly
It cannot tell you a ship is military. Across 326 AIS-matched vessels the classifier predicted military twenty times and was right zero times, and every one of those was in commercial traffic. The cause is enumeration rather than evidence: sixteen military classes compete against the merchant ones, and in a strait full of tankers a 155 by 20 metre contact fits a destroyer about as well as it fits a feeder container ship.
The physics does not help either, and this is the part people find surprising. A carrier flight deck is a flat plate, and a flat plate is a mirror — it reflects the beam away rather than back. Twelve passes over the carrier piers at Norfolk turned up no contact wider than 32 metres. Warships are not findable by being big.
So the tool grew a switch instead of a better guess. "Commercial water" is the default and pushes the military classes down to a twentieth of their weight; "do not ask" removes them from the competition entirely and says so in the output. A naval-approach prior exists for the places it is justified, and even there it is muted whenever the measurements alone already identify the contact — a prior that can overturn a measurement stops being evidence and becomes a way to find what you went looking for. When it does change the answer, the contact says so.
Cross-pass persistence
Run the same area over several passes and match detections by geodesic distance. Anything that repeats at the same coordinates is infrastructure, not a ship.
The first threshold was "present in 70% of passes equals fixed structure", which is backwards. The Long Beach breakwater only fires in two or three passes out of five, because whether a given rock clears the detector depends on speckle and look direction. Worse, adding a sixth, seventh and eighth pass raised the bar from four repeats to six, so the same area went from three fixed verdicts to one while the evidence grew.
The fix was to stop tuning the threshold and compute the null hypothesis from the data. With D detections spread over area A, the chance a pass drops one inside a disc of radius r is the Poisson probability of at least one hit in that disc, which at Long Beach works out to 0.46% per pass.
So 35 repeats against a chance expectation of 2.5 is the finding. The verdict label is just convenience. It also self-adjusts as a harbour gets more crowded, which no constant can.
Mixing ascending and descending passes is a feature, not sloppiness. A sidelobe artifact is tied to look direction and cannot repeat across both. A rock can.
What I would tell you before you start
Build the validation harness early on anything measurement shaped. Every bug worth finding here was a silent numerical error that produced a map which looked entirely plausible. And when you fit a correction, fit both models and score the residuals. The additive "fixed pixel skirt" felt obviously right physically and lost to the multiplicative model on the data, 7.1 metres of error against 4.9.

Comments