Neuromeka and a KAIST team led by Professor Kim Min-jun said on October 2, 2026 that their Spatially Conditioned Diffusion Policy (SCDP) has been accepted at CoRL 2026, the Conference on Robot Learning, held November 9–12 in Austin, Texas. SCDP learns precise manipulation from a single RGB camera, without the wrist-mounted cameras most imitation-learning setups now rely on. Fewer cameras means less wiring and hardware at the hand, which matters most on humanoids.
Key Facts
- CoRL 2026 acceptance announced October 2, 2026; paper on arXiv (2606.14535) since June 12, 2026
- Meta-World Hard (6 tasks, 20 demonstrations): 82.5% average success vs 74.2% for Diffusion Policy with an added wrist camera
- Real cup-handle grasp with 4–7 unseen distractor objects: 85% vs 30% for Diffusion Policy, over 20 trials
- Neuromeka also showed bimanual paper-cup separation on its EIR humanoid using only the head camera
How does SCDP work without a wrist camera?
Wrist cameras earn their place because they see contact up close. SCDP replaces that with attention. According to the arXiv paper, by Seoyoon Kim, Kanghyun Kim, Dongwoo Ko, Yeong Jin Heo and Min Jun Kim, the policy uses predicted end-effector trajectories as “visual attention anchors”. A visual encoder builds feature maps at several scales, and a spatial conditioning module samples features at points along the intermediate trajectory inside the diffusion loop. In plain terms, the policy looks closely at where the gripper is about to go.
On six Meta-World Hard tasks, trained on 20 scripted demonstrations across three seeds, SCDP averaged 82.5%, against 74.2% for Diffusion Policy with an extra wrist camera and 60.3% for DP3 with depth. Real-world tests used a Franka Panda arm and one fixed Intel RealSense camera, with 20 trials per task. In cup-handle grasping SCDP scored 95% in the clean setting and 85% with distractors, against 50% and 30% for Diffusion Policy. Its average across the real tasks was 73.8%, against 50.3%.
What are the limits?
Precision is still the weak point. In USB insertion, SCDP fully seated the connector in 30% of trials; every baseline managed 0%. The authors say the method assumes the useful visual information lies near the end-effector path, which may not hold for tool use, and that heavily occluded scenes remain hard. The EIR humanoid demonstration is a company showcase, not one of the paper’s reported benchmarks.
For Neuromeka, a KOSDAQ-listed cobot maker now building a contract robot foundry in Pohang, the result fits a cost argument. CTO Heo Young-jin, a co-author, said the work links hardware simplification with precision manipulation learning, Newspim reports. Many of the large robot foundation models are trained on multi-camera data that includes wrist views. A method that matches a wrist-camera baseline from one global view lowers the cost of collecting that data and of the robot that runs it.
Frequently Asked
What is SCDP?
Spatially Conditioned Diffusion Policy, a visuomotor policy from Neuromeka and KAIST that uses the robot’s predicted end-effector trajectory to decide where to look in a single RGB image, removing the need for a wrist camera.
How well does SCDP perform?
On six Meta-World Hard tasks with 20 demonstrations it averaged 82.5%, against 74.2% for Diffusion Policy with an extra wrist camera. In real cup-handle grasping with distractor objects it succeeded 85% of the time, against 30% for Diffusion Policy.
Where will the SCDP paper be presented?
At the Conference on Robot Learning (CoRL 2026), held November 9–12, 2026 in Austin, Texas.
Sources & Further Reading
- arXiv — Spatially Conditioned Diffusion Policy: Learning Precise and Robust Manipulation with a Single RGB Camera (Jun 12, 2026)
- Robot Newspaper — Neuromeka and KAIST enable humanoid manipulation without wrist cameras (Oct 2, 2026, Korean)
- Newspim — Neuromeka develops precise robot control technology with KAIST (Oct 2, 2026, Korean)
- CoRL 2026 — Conference on Robot Learning, November 9–12, Austin
- Embodied Wire — Neuromeka plans ₩180bn contract robot foundry in Pohang
- Embodied Wire — Robot foundation models compared