An overhead arm’s-length selfie, specified to the degree — and I mean that literally. Copy it whole.
Needs a reference image. It opens with use the attached image as the main facial reference with high precision and spends a paragraph on preserving an exact identity. Without one, “the subject” is whoever the model feels like.
What it got: nearly all of the concrete half. The camera is up and tilted down, the leopard cap with its curved brim is there, the oversized charcoal graphic tee, layered silver necklaces, the slim watch, the pointed black ankle boots, a translucent takeaway drink, and a whole frame of grey stone paving. Warm late-afternoon light from one side. It is a good result.
The degrees are decoration. This prompt asks for hips rotated 12°, shoulders counter-rotated 8°, head tilted up 28° and turned 10°, the phone 70–85 cm away at roughly 60°. You cannot check any of that in the output, and neither can the model — there is no protractor in the loop. What those numbers really say is “turn slightly, tilt up a bit, hold it high”, and that is what came back. They are not wrong, they are just doing the work of ordinary adjectives while looking like measurements. Compare it with the parts that are checkable — cap, tee, boots, drink, paving — which all landed exactly.
One thing it did not get: the opposite arm extends comfortably away from the body. The drink hand is tucked in close, so the “dynamic triangular arrangement” the last paragraph asks for never forms. Arm placement is a behaviour instruction, and those are the ones that slide.
And the ratio, which is the interesting part. This one ends with --ar 4:5 and came back 928×1152, exactly 4:5. Every prompt here that asked for a ratio in prose at the end of a long block was ignored. The difference is that --ar is not prose — it is a flag the tool parses before the text ever reaches the model. Written as a sentence, a ratio is a suggestion. Written as a flag, it is a setting.