I did not take this photograph, and neither did anybody else. There are two of them here — the one at the top of this post and the one further down — and they came out of the same prompt, run twice, about a minute apart.
I have started keeping the prompts that produce something worth looking at. This is the first of them, in full:
Ultra-realistic candid outdoor photograph of a young adult East Asian woman sitting in a clear shallow mountain stream on a bright summer day. She has straight dark brown-black hair that is naturally wet and slightly clinging to her face, with soft wispy bangs and loose strands framing her cheeks. She looks directly toward the camera with a gentle, calm expression and natural minimal makeup. She is wearing a loose light-gray oversized graphic sweatshirt and light-colored casual shorts, both visibly wet from the clear flowing water. She sits naturally in the shallow stream with the water reaching around her waist and torso, creating realistic ripples and transparent reflections over her clothing. Sunlight filters through dense green forest foliage, producing sparkling highlights across the water and soft dappled light on her face. Large smooth rocks and lush vegetation surround the stream in the background. The image feels like an authentic spontaneous travel photograph taken on a smartphone, slightly low camera angle positioned close to the water surface, natural wide-angle perspective, realistic wet hair and fabric textures, sparkling water droplets, subtle lens flare, soft highlights, shallow depth of field, vibrant but natural greens and blues, highly detailed photorealistic skin, realistic proportions, 4K photography, peaceful summer nature atmosphere, no studio lighting.
What strikes me reading it back is how little of it describes a subject. “A young adult East Asian woman sitting in a clear shallow mountain stream” is the whole idea, and it is over in one clause. Nearly everything after that is an argument with the model’s defaults: authentic spontaneous travel photograph, taken on a smartphone, slightly low camera angle positioned close to the water surface, no studio lighting. Left to itself an image model will light a face beautifully, centre it, and hand you a magazine cover, because that is what most of the photographs it learned from are. Most of the words in this prompt are spent talking it out of that.

Running it again does not give you a variation on the first image. It gives you a different photograph of the same afternoon. The camera has moved back and up. The sky, which is a good third of the first frame, is gone entirely. And she is shading her eyes with one hand — a gesture the prompt never asks for, and probably the most convincing thing in either picture.
What survived both runs is the concrete stuff. Wet hair actually clinging to her face rather than draped for effect. The oversized grey tee, and the fact that it is visibly soaked through. Water at the waist with the light coming up through it. Dappled light, smooth rocks, that particular green. Those are the clauses with real nouns in them, and they are the ones that hold. The vaguer the instruction, the less likely it is to survive a second roll — which is a decent argument for writing these long even when it feels excessive.
The tell is in both, in the same place: the graphic printed on the shirt. It reads as words at a glance and falls apart the second you actually look at it — RMACHEMU, or something equally confident and meaningless. Faces are solved. Hands are solved. Water, which used to be the giveaway, is solved. Text on fabric is not, and it is the first place I look now.