详细信息
Robust human pose generation via uncertainty aware structural reward modeling ( EI收录)
文献类型:期刊文献
英文题名:Robust human pose generation via uncertainty aware structural reward modeling
作者:Zhang, Deliang[1]; Dai, Songyin[1]; Xu, Bingxin[1]; Liu, Hongzhe[1]; Pan, Weiguo[1]; Xu, Cheng[1]
第一作者:Zhang, Deliang
机构:[1] Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing, China
第一机构:北京联合大学北京市信息服务工程重点实验室
通讯机构:[1]Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing, China|[11417103]北京联合大学北京市信息服务工程重点实验室;[11417]北京联合大学;
年份:2026
卷号:6
期号:1
外文期刊名:Discover Artificial Intelligence
收录:EI(收录号:20263221256834);Scopus(收录号:2-s2.0-105046569325)
语种:英文
外文关键词:Aberrations - Diffusion in liquids - Human form models - Image processing - Musculoskeletal system - Semantics - Topology
摘要:Conditional diffusion models offer a versatile paradigm for controllable image synthesis, yet faithfully adhering to intricate spatial constraints, such as human pose skeletons, continues to pose significant challenges. Despite progress in control architectures and reward-guided fine-tuning, outputs frequently exhibit structural aberrations, joint misalignments, and anatomically infeasible configurations. We contend that these shortcomings arise primarily from two core deficiencies: pixel-based rewards inadequately encapsulate the perceptual and topological nuances of skeletal forms, and reward signals grow unreliable amid diverse or out-of-distribution samples. To address this, we introduce Dream Your Pose, a perceptually informed framework for pose-conditioned generation that prioritizes structural fidelity. Our method incorporates a multi-channel, structure-sensitive reward mechanism, harnessing perceptual features like local contrast, edge gradients, and spatial continuity to more accurately gauge pose congruence. Critically, we integrate an uncertainty-aware regularization paradigm—drawing from principles of uncertainty modeling in learning—to adaptively modulate reward influence, thereby mitigating the effects of spurious or ambiguous feedback and fostering robust training dynamics. Rigorous evaluations on the OpenPose-ControlNet dataset reveal substantial gains, compared with the retrained ControlNet baseline, including a 25.3% relative uplift in Object Keypoint Similarity (OKS) and a 15.6% enhancement in Probability of Correct Keypoint at 0.5 (PCK@0.5), underscoring improved keypoint precision and holistic skeletal integrity. These advancements yield images with superior visual coherence and anatomical plausibility, without compromising semantic fidelity or perceptual quality. ? The Author(s) 2026.
参考文献:
正在载入数据...
