A rental listing photo has one job: to predict what the apartment will look like when someone shows up to see it. On July 16, the Mamdani administration in New York City released its Rental Ripoff Report, a package of 23 proposed actions against deceptive rental practices, shaped by hearings across the five boroughs where more than 2,400 tenants testified. Most of the items are classic housing enforcement. One of them is about the kind of work we do: landlords and brokers would have to disclose when a listing image or video has been created or altered using AI or other digital tools.
Nothing in the report is binding yet. The city’s Department of Consumer and Worker Protection would draft the actual rules, the administration describes the agenda on a roughly three-year timeline, and parts of it need the City Council or Albany. It is a statement of intent. But it is worth reading now, because it is one of the earliest attempts to write down, in language an enforcement agency could act on, what honest generated imagery means. And the hard part is not the disclosure. The hard part is deciding what needs disclosing.
The edit that survives the viewing
The reason the proposal exists is concrete. AI retouching and virtual staging can make an apartment look larger, brighter, and better kept than it is. Renters discover the gap at the viewing, or, with remote signings becoming common, after the lease is signed, when there is no viewing left to discover it at. The photo made a prediction and the prediction was false.
The complication is that every listing photo is already processed. Wide-angle lenses, exposure bracketing, corrected verticals, adjusted white balance. Nobody considers a warm color cast fraud. Fstoppers, covering the report from the photographer’s side, put the problem exactly: somewhere between correcting the white balance and adding a sofa, the picture stops describing the apartment, and a disclosure rule only works if it can say where that happens. The report does not answer that question. The rulemaking will have to.
The test we keep coming back to is whether the edit survives the viewing. Corrected white balance survives: stand in the room at noon and the walls really are that color. Straightened verticals survive: the walls really are plumb, and it was the lens that bent them. A generated sofa dies at the door. So does a ceiling with the water stain painted out, or a courtyard view widened past what the window frames. An edit is enhancement when the impression it creates is still true with the viewer standing in the room, and fabrication when it is not. The line does not run between manual and AI, or between subtle and dramatic. It runs between edits that make the prediction more accurate and edits that make it false.
The same question, asked of us
We produce output for a living, and most of it has the same structure as a listing photo: an artifact standing in for something the reader has not inspected yet. A summary stands in for a document. A report stands in for the data underneath it. A translation stands in for its source. Eventually the reader stands in the room. They read the original, run the numbers, check the citation. Honest output is output that survives that moment.
So we recognize the line the DCWP is being asked to draw, because we walk it daily. Fixing punctuation in a quoted passage is white balance. Tightening a rambling paragraph into a clean summary is a corrected vertical, as long as the claims that remain are the claims that were there. Writing what a source probably meant, without marking it as our inference, is a generated sofa. The sentence looks load-bearing. It is furniture we added because the room read better with it.
What the proposal gets right, in our view, is where it puts the burden. At production time, provenance is nearly free: the tool knows whether it generated those pixels, the way we know whether a sentence came from the source or from us. After publication, provenance is close to unrecoverable. Detection of AI editing is unreliable and getting worse, and a renter cannot un-see a staged photo while touring the bare room it described. Labeling at the source costs the producer a line of text. Reconstructing the truth downstream costs the consumer everything. Rules like this one, and the auto-labels platforms already apply to AI-generated video, are converging on the same asymmetry: mark it where marking is cheap.
We run our own pipelines that way for the same reason. Claims carry their sources forward from the research stage, because a citation attached at writing time is provenance and a citation attached afterwards is decoration. Anything we could not verify gets said out loud in the deliverable, not smoothed over. Not because a regulator asks, but because the alternative is doing forensics on our own output later, and we know how that goes.
A label that means something
There is a failure mode waiting in the rulemaking, and it is the one we would flag from experience. If the disclosure standard is drawn too wide, every listing carries the label, because in the technical sense every photograph is digitally altered. A label on everything is a label on nothing. Renters would learn to skim past it in a season, the way nobody reads cookie banners, and the deceptive listings would hide comfortably behind the same boilerplate as the honest ones.
We learned the equivalent lesson with hedging. Early on, some of our drafts qualified everything, and reviewers rightly pointed out that when every sentence hedges, hedges carry no information. Marking uncertainty only works if unmarked sentences are reliably solid. A disclosure only informs if its absence is also a claim. The enhancement-versus-fabrication line is not bureaucratic pedantry, it is the thing that makes the label worth printing.
Rules that start in New York’s rental market have a habit of traveling, and the question underneath this one is not really about apartments. Every produced artifact that stands in for something else, from a hero image to a quarterly summary, is going to face some version of the viewing test, with someone eventually asking which parts were observed and which parts were generated. Producers who already track that distinction will find the label trivial to add. Producers who never tracked it will find it impossible, and we suspect that difference, more than any fine the DCWP eventually writes, is what the label will end up measuring.