05 · Identity, outfit, pose
Person, Garment and Pose
Choose the person, the outfit and the body position as three separate inputs. One reference image supplies the identity, one the complete outfit, and a third supplies the pose. The pose is read from a photo as a skeleton, a stick figure of lines and points, which you can edit.
This page covers two downloadable workflows: a local one, and an optional cloud one that sends the same three photographs to six providers. The local one needs extra custom nodes for pose detection and editing, listed in Getting Started.
- You’ll learn
- How to detect a pose skeleton from a photo, edit it and reuse the edit, how three references are given three roles, and how to read a comparison between providers.

What goes in, what comes out
How the workflow works
Explained earlier:seed, guidance and steps · reference images and written roles (03) · outfit transfer (04) · decode and save

From this workflow on, nodes are named as they appear in ComfyUI.
Load the three references in order
Three LoadImage nodes, top to bottom: identity (face and hair), outfit (the complete look) and pose (a photograph showing the body position you want). The prompt calls them image 1, image 2 and image 3, so keep the order.
Detect, edit and reuse the pose

Group 3b, with the pose detector on the left. - DWPose EstimatorFinds the skeleton in the pose photo.
- Pose Input ToggleSwitches between detecting the pose again and reusing your edited pose.
- OpenPose StudioThe editor. Drag joints here, then press Apply.
- PreviewShows the skeleton that will be used.
The pose is handled in a fixed sequence. Follow it in this order:
- Detect. Run the workflow once. DWPreprocessor analyses the pose photo and produces a skeleton: coloured lines and points for the head, shoulders, arms, hands, hips, legs and feet.
- Inspect. Look at the skeleton preview. Check that both arms and both legs were found and that left and right are not swapped.
- Edit. Open OpenPoseStudio and drag any joint to a new position, for example to move a hand to the hip.
- Apply. Press Apply in the editor so the edited skeleton is stored in the node.
- Select REUSE EDITED POSE. Set PoseInputToggle to REUSE EDITED POSE. If you leave it on detection, the next run detects the pose from the photo again and overwrites your edit.
- Generate again. Press Run. The result now follows the edited skeleton.
Change: one joint. Keep fixed: identity, outfit, prompt and seed. Inspect: whether only the body position changed.
How the pose reaches the model
The skeleton is drawn as a picture and given to the model as a third reference image, in the same way as the identity and outfit photos. Some workflows use a separate add-on model called a ControlNet to force a pose. This one does not. The skeleton is a visual reference that the model is asked to follow, so it guides the body without fixing it exactly. Hands and depth can still differ.
Using the skeleton instead of the pose photo keeps that photo’s clothes, face and background out of the result. The prompt assigns the three roles in writing, as in 03, and describes the pose in words as well. Describing the hands helps where the skeleton is unclear.
Generate locally
The same models and controls as 01A, on an 832 × 1216 image. The workflow uses the recommended 4 steps and CFG 1 with a fixed seed.
The second workflow: six cloud providers
The file B_Cloud_Model_Comparison.json is a separate, optional workflow. Its purpose is to send the same task to six paid image providers in one run, so you can compare how each handles identity, outfit and pose.
- Inputs: the same three photographs (identity, outfit, pose) and one shared written instruction that all six providers receive
- Outputs: six images, one per provider, each saved separately under a provider-specific filename
- Needs: a Comfy account with credits. No local model files. Every full run can charge for six images. The note inside the workflow suggests budgeting roughly 120 to 160 credits per complete comparison. Check the price badges, because prices change
| Provider | Model selected in the workflow | Saved size and quality | How it receives the references |
|---|---|---|---|
Nano Banana Pro (gemini-3-pro-image-preview) | 2:3, 1K | Resized to 768 × 1152 with white borders added, and sent together | |
| OpenAI | GPT Image 2 (gpt-image-2) | 1024 × 1536, medium quality | Original photographs |
| ByteDance | Seedream 5.0 Pro | 832 × 1248 (1K) | Original photographs |
| Black Forest Labs | FLUX.2 Pro | 832 × 1248 | Original photographs |
| xAI | Grok Imagine Image 2.0 | 2:3, 1K, medium quality | Original photographs |
| Kling | 3.0 Omni Image (kling-v3-omni) | 2:3, 1K | Resized to 768 × 1152 with white borders added, and sent together |

Why this is not a controlled ranking
- Input preparation differs. Google and Kling receive resized copies with white borders, because their nodes need all three images to be the same size. The other four receive the original files
- Sizes and quality tiers differ. Each provider is set to its own roughly one-megapixel option, and OpenAI and xAI have a quality setting the others lack
- Seeds are not comparable. Seed 22 is fixed where supported, but a seed means something different in every model, and OpenAI and Kling do not promise repeatable results
- The pose input differs from the local workflow. The cloud providers receive the original pose photograph. The local workflow uses the extracted skeleton
To test one provider only, switch off the other five Save Image nodes before running. Select them and press Ctrl+M. This is called muting a node.
Local and cloud results
These images were saved from earlier runs, and some of the cloud images were made with earlier versions of the input files. Treat them as illustrations of what each route can produce, not as a controlled comparison or a ranking.
Check the result
- Identity: the face matches image 1
- Outfit: the garments keep their proportions, seams, colour and footwear from image 2
- Pose: the joints follow the skeleton, including hands and crossing legs, and not the clothes or setting of the pose photo
- Hands touch the clothes convincingly, with no crushed fabric or hands merging into a hip
- For the cloud comparison: note the charge and the time for each provider alongside the image
A face belongs to someoneUse an identity photo only with that person’s permission for this purpose. Use it responsibly.
Download the workflow
- A_Person_Garment_Pose.jsonLocal workflow with pose detection and editing
- B_Cloud_Model_Comparison.jsonSix-provider cloud comparison · uses Comfy account credits
- HERO_model_identity.png, L06C_dressed_anchor.png, Pose Reference Prompt.jpgThe example identity, outfit and pose images. Put them in ComfyUI’s input folder, or load your own
Click a filename to download it, or right-click it and choose Save link as. Keep the .json ending. Then drag the file onto the ComfyUI canvas, or use Workflow → Open.











