About Me
- I am Zhang Jinquan(张锦权), an Artificial Intelligence master's student at Shenzhen University and a member of the World Model Group at the Guangdong Laboratory of Artificial Intelligence and Digital Economy (Shenzhen), led by Tian Qi (田奇), Chief Scientist of Huawei Terminal BG. I received my bachelor's degree from Jiangsu University.
- My research focuses on robot learning and real-world robustness of VLA / WAM models, especially post-training and fine-tuning for physical environments.
- My work has been accepted by IROS and AHSWN, and I have filed multiple invention patents.
- I am currently interning at Roboscience, working on VLA / WAM fine-tuning and post-training for strong simulation performance and physical-world robustness, including robot evaluations in large shopping-mall and supermarket shelf scenarios.
I hope to join a dynamic and goal-driven robotics team. If you are interested, please email me at 2410815010@mails.szu.edu.cn.
News
- 2026.06: My paper Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking was accepted by IROS 2026.
- 2026.05: My paper VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks was accepted by the journal AHSWN.
- 2026.05: My invention patent 基于语义嵌入和注意力双一致性的机器人控制方法、系统、终端及存储介质 was granted. Patent No.: ZL 2026 1 0211734.7.
- 2026.05: My invention patent 一种基于视觉语言动作模型的机器人控制方法、系统、终端及存储介质 was granted. Patent No.: ZL 2026 1 0219615.6.
- 2026.05: My invention patent 一种基于视觉语言模型的机器人操作方法、系统及终端 was granted. Patent No.: ZL 2026 1 0200293.0.
Achievements
Papers
- Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking. Accepted by IROS 2026.
- VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks. Accepted by AHSWN.
Patents
- 基于语义嵌入和注意力双一致性的机器人控制方法、系统、终端及存储介质. Granted. Patent No.: ZL 2026 1 0211734.7.
- 一种基于视觉语言动作模型的机器人控制方法、系统、终端及存储介质. Granted. Patent No.: ZL 2026 1 0219615.6.
- 一种基于视觉语言模型的机器人操作方法、系统及终端. Granted. Patent No.: ZL 2026 1 0200293.0.
- 一种约束驱动轨迹生成方法、系统、终端及存储介质. Status: under substantive examination.
- 一种柔性压力传感器及其制备方法. Status: under substantive examination.
- 一种基于未来潜在表征一致性约束的视觉语言动作模型鲁棒训练方法. Status: under preliminary examination.
Selected Publications
Robot Learning and VLA Robustness
Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking
Zhang Jinquan(张锦权) et al.
- Identifies policy-critical action-to-vision attention hijacking, where physical adversarial patches divert action-conditioned attention away from task-relevant regions.
- Proposes AGSD, an EOT-optimized printable patch that jointly hijacks policy-critical attention and disrupts vision-language semantic alignment.
- Introduces SARF, a zero-inference-overhead defense that fine-tunes only the visual encoder using feature anchoring, attention correction, and language-guided geometric consistency.
- Reduces OpenVLA failure under AGSD from 100% to 17.0-56.8% on LIBERO and improves real PiPER success from 9.7% to 61.7%.
VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks
Zhang Jinquan(张锦权) et al.
- Studies VLA robots as mobile edge nodes in WSN-assisted smart environments and evaluates their robustness against physical adversarial threats.
- Introduces VASA, a printable-patch stress test that distracts action-conditioned cross-attention through policy-critical attention hijacking.
- Proposes APFT, a zero-inference-overhead defense that stabilizes spatial-temporal attention and enforces geometric consistency.
- Reduces OpenVLA failure from 100.0% to 25.9% in LIBERO and improves real-world success from 23.0% to 67.4% across 2,000 severe patch-attack trials.
Selected Projects
Real-Robot Learning Systems
Aloha Pi0.5 LeRobot: A Real-Robot Learning Pipeline
Zhang Jinquan(张锦权)
- Builds an end-to-end pipeline for AlohaMini teleoperation, LeRobot-format datasets, dataset cleaning, Pi0.5 training, and 4090-to-Orin remote inference.
- Separates real-time robot IO and safety on Orin from heavy VLA policy inference on a remote GPU server.
- Provides public engineering documentation for reproducible real-robot learning without exposing private datasets, paths, IPs, or hardware identifiers.
Research Interests
- Robot Learning: data-driven manipulation, real-robot systems, and closed-loop deployment.
- VLA: Vision-Language-Action policies for general robotic manipulation.
- WAM: World Action Models for action-aware embodied prediction and planning.
- Post-Training: adapting pretrained robot policies for downstream tasks and robustness.
- Fine-Tuning: structure-aware and parameter-efficient adaptation for embodied models.
Experiences
Demo Videos
Contact
- Email: 2410815010@mails.szu.edu.cn
- GitHub: https://github.com/jjjjqqqz101
- Homepage: https://jjjjqqqz101.github.io/