Projects

Selected robotics, embodied intelligence, and visual perception projects.

FluxVLA

A one-stop VLA engineering platform for embodied intelligence

Contributor Website Code

TL;DR

FluxVLA Engine is a full-stack, end-to-end engineering platform for embodied intelligence. It provides a unified workflow spanning data preparation, model training, simulation evaluation, and real-robot deployment.

My Role

Open-source contributor focused on model and simulation capabilities. I integrated FastWAM’s uncond, joint, and IDM modeling paradigms together with Wan2.2, MoT, and ActionDiT components. I also connected data preprocessing, distributed training, and LIBERO evaluation, and integrated LIBERO-Plus into the platform.

Highlights

  • Model integration. Adds FastWAM’s uncond, joint, and IDM paradigms together with Wan2.2, MoT, and ActionDiT components.
  • End-to-end training and evaluation. Connects data preprocessing, distributed training, and LIBERO evaluation in a unified workflow.
  • LIBERO-Plus support. Enables zero-shot evaluation of standard LIBERO checkpoints, multi-task rollouts, and aggregated result reporting.

Visit the FluxVLA website for documentation and setup guides, or browse the source code on GitHub.

FluxDAgger

A Model-Decoupled DAgger Pipeline for Dual-Arm Robotic Manipulation

Rollout to takeover demo frame
Rollout → Takeover. A single DAgger episode with autonomous-to-human handover at 0:30.
FluxDAgger system architecture
Model-decoupled architecture. Policy and reward modules plug in via standardized ROS topics.
Baseline data collection system
Baseline collection system. Cameras + master arm (demonstration) + slave arm (execution) on the AgileX Piper platform.
End-to-end data flow
End-to-end data flow. Raw multi-modal capture → Parquet → MP4 / NumPy / takeover clips / reward annotations.
Timestamp synchronization
Timestamp synchronization. Per-camera image buffers matched to robot states inside a bounded sync window.
Original CAN-bus topology
Hardware — before. Original shared CAN-bus topology of the AgileX Piper four-arm platform.
Modified CAN-bus topology
Hardware — after. Per-arm dedicated USB-CAN interfaces enabling independent enable/mode/control.

TL;DR

Deploying DAgger on a real dual-arm robot is far more than a policy-inference problem — it tangles policy, hardware, teleoperation, multi-camera sync, and post-processing into a single brittle stack. FluxDAgger decouples these concerns behind a small set of ROS topics, so swapping the VLA model or the reward model never touches the collection logic.

Contributions

  • Model-decoupled architecture. Policy inference lives in an external project/service; the collector consumes a fixed action topic, so one DAgger workflow serves arbitrary VLA policies.
  • Human-in-the-loop DAgger loop. Autonomous rollout, online human takeover, and per-frame source tagging (rollout vs. correction) within a single episode, with multi-camera + joint-state timestamp alignment.
  • Reward-pluggable infrastructure. A Qwen3-VL reward module is integrated via an independent interface for both online ROS publishing and offline batch annotation, enabling reward-guided dataset filtering.

Method & Hardware

FluxDAgger is organized as a set of ROS Noetic nodes — camera, sync-observation, model-inference, DAgger controller, DAgger collector, and reward node — communicating through standardized topic interfaces. The policy and reward models become drop-in modules rather than first-class citizens of the collector. On the hardware side, the stock AgileX Piper platform shares one CAN bus across arms; FluxDAgger introduces a per-arm USB-CAN topology so each arm’s enable state, mode, and commands can be managed independently.

For interactive demo videos and the full system walkthrough, visit the project page.

FluxThemis

A standalone VLA evaluation framework for multiple models and simulators

Contributor Code

TL;DR

FluxThemis is a standalone VLA evaluation framework that supports multiple models and simulators. It separates model inference from simulation evaluation and isolates simulator dependencies, making evaluation pipelines easier to reproduce and extend.

My Role

I led the development of the LIBERO-Plus evaluation pipeline. ROS decouples model inference from simulation evaluation, while separate Conda environments isolate LIBERO, LIBERO-Plus, and robosuite dependencies.

Highlights

  • Large-scale robustness evaluation. Covers 10,030 LIBERO-Plus robustness tasks.
  • Flexible analysis. Supports metadata filtering and grouped success-rate statistics.
  • Parallel execution. Provides multi-process rollouts for faster evaluation.

Browse the source code on GitHub.

Road Detection in Semi-Structured Outdoor Environments

The SUNSET benchmark and TEAR method for end-to-end road and boundary detection

Research Project Paper

Research Problem

This project studies end-to-end detection of roads and their boundary lines in semi-structured outdoor environments that lack artificial lane markings but retain visible road trajectories.

Contributions

  • Benchmark construction. Selected semi-structured scenes from the public DeepScene, YCOR, and GOOSE datasets. Using an enhanced Labelme curve annotation tool, we labeled left and right road boundaries with high-order Bézier curves and paired them with road instances to build the SUNSET dataset and benchmark.
  • Method design. Contributed to TEAR, which uses interchangeable dual instance decoders to disentangle road and boundary-line instances and hierarchical bipartite matching to associate them.
  • Research output. TEAR of the SUNSET introduces the benchmark and end-to-end method and was accepted by IEEE ICME 2026 (CCF B).

Gesture Recognition on Horizon X3 Pi

Efficient gesture recognition and mobile-robot deployment (undergraduate thesis)

Undergraduate Thesis Edge AI YOLOv5s ROS2
Real-time gesture recognition. YOLOv5s on Horizon X3 Pi after optimization — 30 FPS detection with 0.36% mAP loss versus the GPU baseline.
Gesture-controlled mobile robot. Recognized gestures mapped to ROS2 motion commands driving the platform in real time.
Human tracking. Detection-driven person tracking module integrated with the mobile-robot motion-control node.

TL;DR

An undergraduate-thesis project on deploying efficient gesture recognition on the Horizon X3 Pi edge board and integrating it into a mobile-robot system. The work covers model selection, edge-side optimization, and ROS2 integration with simulation and real-robot validation.

Contributions

  • Edge-deployable detector. Trained YOLOv5s on the HaGRID static-gesture dataset, reaching 86.60% mAP, and benchmarked against the YOLO and DETR families.
  • Edge optimization. Optimized and deployed YOLOv5s on the Horizon X3 Pi, improving inference from 20 FPS to 30 FPS with only 0.36% mAP degradation.
  • Robot integration. Built ROS2 motion-control nodes and integrated gesture recognition with human tracking on the lab’s mobile robot.

Method & Platform

The pipeline combines a YOLOv5s detector (HaGRID-trained) with Horizon-toolchain-based quantization and graph optimization for the X3 Pi BPU. Recognized gestures are mapped to discrete motion primitives published as ROS2 Twist messages; a parallel human-tracking module shares the detection backbone to drive the platform’s follow behavior.