Listen to this content
A science and industry cooperative project is conducting multi-modal sensor capture missions for different scenarios and in diverse environments. This dataset aids R&D for robotics and autonomy applications and includes an innovative approach to providing a millimeter-precision, ground truth foundation.
A multimodal and open training robotics database, based on diverse “missions” captured with a quadruped robot, the first of its kind, has been created by researchers at ETH Zürich (Eidgenössische Technische Hochschule Zürich; Federal Institute of Technology Zurich). Working in cooperation with multiple geospatial and geomatics manufacturers and solution providers, including Leica Geosystems (part of Hexagon), the ETH GrandTour team built a multimodal capture system. They then conducted 49 initial data-capture missions across various environments and scenarios in which such robots are envisaged for use.
There are existing training databases for other robotics and autonomous platform classes, including aerial, wheeled, motor vehicles, handheld and UGV. For example: KITTI for autonomous driving, EuRoC for small UAV, and TUM VI for visual-inertial odometry. But no comprehensive dataset has been compiled for legged robots, particularly quadrupeds. This project and dataset have been named “GrandTour” (paper) and have been made openly available via ETH and the collaborative developer platform, GitHub.
Learning Challenges
The preeminent approach for robot spatial sensing to inform navigation and locomotion, as envisaged for autonomous applications in particular, is state estimation and simultaneous localization and mapping (SLAM). SLAM is a multiple synchronized set of sensors for real-time modeling of the environment (exteroception)and the robot’s internal feedback from its state (proprioception).
Capturing real-time data to provide action-perception feedback can be much more difficult for quadruped robots than other platforms. For example, there can be motion-induced disturbances, like slipping on a surface, unstructured and harsh environments (requiring multiple sensor types, which adds to complexity), rapid or sudden changes in locomotion profiles, and the need for flexible algorithms for diverse, sometimes harsh environments (e.g., urban, forest, snow, smoke).
In humans, this meld of proprioception and exteroception combines innate perception, instinct and learning through real-world experience. An infant immediately begins to make the crucial connections between how they perceive their surroundings and action-perception feedback bindings are gathered and refined over time. This is essentially the approach taken by robotics R&D, except that, as in the case of GrandTour and other datasets, these experiences can be preloaded into the bot before it embarks on its journeys.
What the research team developed was a multimodal payload, dubbed “Boxi,” for the ANYmal quadruped platform used for the missions. The resultant data, though, is not limited in its utility to this one platform; it could be applied to others.
The Sensor Stack — Boxi
Boxi is, to put it simply, a powerful beast of a reality-capture (RC) payload, yet quite compact. The GrandTour team sought to include a wide variety of representative modalities that have proven beneficial for robotics exteroception. Turcan Tuna, a Ph.D. candidate at ETH Zurich and co-lead researcher on the project, said they chose off-the-shelf sensors in the same size, weight, and cost classes that downstream users of the training data would potentially use for their robotics and autonomy applications.

The proprioception sensors standard to the ANYmal-D quadruped include joint encoders and a proprietary inertial measurement unit (IMU). An additional seven IMUs have been added, including some that are integrated into other sensors (e.g., lidars, cameras, GNSS and a custom Leica AP20). For exteroception, Boxi also packs two lidars, 10 cameras, an RGB-D camera in addition to the lidar and six depth-cameras that are available on ANYmal-D quadruped. In addition to these, a dual-antenna GNSS (Novatel CPT7), and a Leica MS60 in combination with the custom AP20 smart pole for ground truth (more on that later).
“During the GrandTour project, the main purpose was the dataset collection,” Tuna said. “But through our collaboration with Leica Geosystems, autonomy became a secondary mission as we built a solution that works at the click of a button.” Downstream researchers have also been able to use the rich training data to aid their development of autonomous solutions.
In evaluating raw captured data already available, the GrandTour team came to some interesting conclusions. One was that there was little data combining both proprioception and exteroception and, of what was already available, it was often from limited combinations of sensors. In addition, evaluation against ground truth is often limited to data from the reality-capture sensors themselves (not necessarily designed for, or capable of, higher precision).
Another key finding is that some of the research community are not necessarily aware of what is out there, and the implications of the precision and accuracy that components and combinations thereof might bring.
“There needs to be greater awareness in the robotics sector of the technologies already available and their high accuracy and reliability. If the solutions are not readily available to the robotics sector, they cannot fully imagine how they can leverage these systems,” Tuna said. “An example of this is the second GNSS antenna, which can resolve your heading. This came about because some of my colleagues and developers were really hungry for this data. They might be doing navigation, and so don’t need to know exactly where their robot is, as they don’t want to rely on coordinates, but they really want to know where the robot is heading.”
Discovering wider selections and combinations of sensors and sensor data begets a hunger for even richer data. “There are aspects they don’t necessarily care about, initially,” said Tuna. “They might be OK with 2 m of accuracy in position, but they really want to know if the robot is oriented a certain way. When we walk around, we have our heads down, looking at the phone in front of us to see which way we are oriented on the map. In a way, it should be the same in robotics.”
So much of the success of GrandTour has been a “build it, and they will come” proposition. “Customers may not know what they want until they see it, and that applies even to research. No one came to us and asked for a particular type of solution for a particular application. Once we started putting the data into their hands, however, that was when they became excited,” Tuna said.
The variety of sensors also supports work in the field of kinematic inertial automatic estimation, which is another example of how robotics and reality-capture technologies mimic human-like attributes. Basically, this is state estimation using the leg joints and inertial measurements from the IMU. Developers have discovered that kinematic measurement estimation is a very good prerequisite and a strong initial guess for lidar and camera operations. This is because such sensors do not have the sampling rates of, say, the 400 Hz of joint encoders. GNSS might only be updating at 20 Hz. Boxi has IMUs that range from 45 Hz to 400 Hz. By estimating a state, a sensor with a higher sampling rate can better resolve the 3D spatial position and orientation of its captured data.
Synchronization was therefore crucial to the integration of so many sensors. The researchers of GrandTour discovered that while GNSS can provide precise time, many GNSS boards are designed to be time servers, not time slaves. Their default behavior when there is no GNSS clock remains undefined, and much of it is unexplored. For example, when they lose sky view, they might reset to a default time (e.g., Jan.1, 1970). This was the case for the research team, but they were able to use the Novatel Span CPT7 as a time master by sending an approximate time from the local machinery to the device.
The ability to capture in mixed environments is a function of the capability of individual sensors. Light profiles, particulate matter in the air (i.e., smoke or mist), surface reflectivity, glare, etc., all vary, even within classes of lidar and cameras. Performance in smoke-filled locations is always a question, especially when discussing common scenarios where a robot would be extremely useful, such as search and rescue.
Lasers, even the high-end lasers used in top-tier surveying total stations and high-end lidars, may be fine in smoke for up to 10 m to 20 m, depending on the smoke density, but may struggle beyond that as they are based on visible light. As a rule of thumb, you can expect a laser to work for distances you can also see with your eyes. With very dense smoke, a laser might fail at 3 m. Radar is often floated as the solution, having an advantage in this case over other sensors, but also disadvantages compared with others. Again, the strengths and weaknesses of specific sensor classes are why multimodal stacks are common in robotics and autonomy.
Boxi did not include radar, but the GrandTour research team is exploring that. “I was really trying to get a radar that is small enough, but powerful enough to be useful,” said Tuna. “In the field of radar sensors, companies are either going for an autonomous driving solution, the clear business case, and some of them are going for larger-scale applications, for example defense purposes. I really wanted a solution that did not just give me the peaks of the Doppler measurements, not just a point cloud, but rather the actual raw measurements and configurations. According to the experts we’ve consulted in the field, this is where the real benefit of radar comes in: it is so versatile, and you can extract so much information out of it. The peak measurements are just one type of data.”
Boxi did not need extraordinary processing resources. “Luckily, we already had enough computing power: Nvidia AGX ORIN and an Intel i7 NUC CPU,” said Tuna. “So, we have two compute units that can manage the data live. We are currently using these systems to process and compress the data live to be able to store substantial amounts of data from, for example, 10 cameras and multiple lidars.”
He added, “It allows us to support autonomy as well, in the sense that we can do light-dimensional automatic SLAM, semantic segmentation, and object tracking via machine learning. We wanted to dedicate our compute units to certain purposes, but our analysis suggests that you can actually run everything on a single GPU unit; you just need to be a bit careful about how you manage your data and on-board computing.” Autonomy solutions developed with the aid of this training data might well include neural processing units (NPUs), but for data capture, simpler, more affordable processors may suffice. This is why edge computing is having a big impact on the reality-capture sensor sector.
Getting to Ground Truth
A particular robotic application might not necessarily require high precision, and that may be true. But what precision are you truly getting? All too often, you will see instruments, devices, components, and RC systems touted as “centimeter” precision. Is that relative integrity between individual points, or relative to a specific reference frame? It’s instructive to hand a datasheet for, say, an RC handheld SLAM unit to a surveyor or geodesist and ask how they would test its precision and accuracy. It can be an involved process and then consider how you would do this for a multimodal system.
Spatial testing requires baseline references, typically using instruments that represent the highest levels of driftless, field-observable precision and ability to tie to georeferenced frames. Putting aside specialized instruments for industrial and manufacturing metrology, when it comes to geospatial applications, the ubiquitous surveyor’s robotic total station (RTS) is often the final, unimpeachable resource. And it can work in GNSS-denied environments.
For Boxi, the GrandTour team chose to add a laser-target prism, a Leica MS60 with a modified AP20, which was used to stream the total station coordinates to Boxi. The MS60 was set up and leveled on a heavy tripod for each mission (indoors and out). The MS60 tracks the prism and feeds the highly precise positional data back to Boxi, serving as the millimeter absolute position reference for all other spatial data.
With many sensors and solutions, such as real-time kinematic GNSS or optimized bundle adjustments for images, everything can sound like the best option out there, but in fact it might not be as accurate as you think, especially when integrating multiple sensors. Developers in the robotics and autonomy sector might not be aware of the benefits that strong engineering projects and solutions can bring. Why such high precision?
Consider the boom in robotics for construction: autonomous heavy equipment and layout robots. Even in an industrial setting, one might envision a humanoid robot, head down over a workbench, only concerned with the position of its hands and fingers. That is a limited view of the potential for mobile, autonomous robots that can move about the factory, tasked with more than one type of assignment: delivering, working on larger assemblies, inspecting and more. It is no wonder that robotics systems in construction and industrial settings have been implementing RTS and other laser and prism tracking systems because they eliminate a lot of uncertainties.
Boxi, with its RTS component, is capturing data precise enough for all of the above. There were challenges, though. “First, because the total station is only providing the position measurements, we have to build a sensor fusion solution with inertial measurements to get the poses and then control through that,” said Tuna. “Another challenge was the time-consuming manual leveling of the tripod that the MS60 was mounted on. Not insurmountable issues, but for the GrandTour missions, we turned to people with RTS experience to minimize the time spent.”
The MS60 + AP20 combination was used for both indoor and outdoor missions. While GNSS could provide a baseline for spatial reference outdoors, the RTS was used for all scenarios to provide a consistent and highly precise spatial baseline. The Leica Geosystems AP20 tilt prism was the first of its kind, an attachment for a field surveying pole that reconciles the position of the tip of the pole on the ground, no matter the orientation of the prism at the top. “We had to make one modification,” said Benjamin Müller, a principal hardware engineer at Leica Geosystems, who worked with the GrandTour team on this implementation. “We have a pretty good time synchronization over Bluetooth of milliseconds. However, to do this tiered functionality, we streamed the data from the MS60 to the AP20 and then over a ROS node on the AP20 over the existing USB-C to Boxi.”
“This meant the GrandTour team had really good time synchronized data and low latency directly to Boxi, meaning they did not have to process data after every scan,” said Müller. “Our first approach was a separate data logger, and in post-processing we tried to fit the trajectories to each other. Then we came up with a better solution for them: we could create a special, custom version of the AP20 to ensure that they had one dataset recorded at the same time.’”
While the MS60 was used as the RTS for GrandTour last fall, Leica Geosystems released the TS20, which is also AP20 compatible. The TS20 has the distinction of being the first RTS of its kind that has NPU on board. An initial benefit of on-board NPU is enhanced prism target recognition and tracking, which could be beneficial for robotics, as well as surveying.
“The TS20 is also Ethernet capable,” said Müller. “So, among other things, that will enable improved time synchronization. There is also definitely more to come. As we said when the TS20 was announced last year: ‘be ready.’ I can’t say much more at this time, but be ready. The TS20 will become a robotic total station for robotics, not only for surveying.”
The Missions
“While GrandTour is vast, we particularly tried to focus on outdoor environments,” said Tuna. “At the same time, some of the missions bridge indoor and outdoor environments, which allows us to showcase to our users how sensor fusion across different modalities can complement GNSS-based localization. In today’s robots, we expect quadrupeds, humanoids, and other mobile robots to work indoors and outdoors, for example going into tunnels or operating in GNSS-denied scenarios without relying solely on dead reckoning.”
The 49 initial missions include indoor, outdoor and mixed locations. Some with full GNSS, partial GNSS (along some of the route), or completely GNSS-denied environments. Locations and routes were also chosen to represent anticipated use-case scenarios. For example: construction, industrial, and search and rescue.
For more details on each mission, refer to Table 4 and Figure 7 on pages 11 and 12, respectively, of the excellent project paper: GrandTour: A legged robotics dataset in the wild for multi-modal perception and state estimation, Frey, Tuna, et al.
Missions were typically around 300 m in length. “Partially, because of the battery [capacity] of the robot, and partially to provide full coverage of the total station,” Tuna said. Going beyond 300 m would require extra setups and battery swaps, but this was not seen as a problem, as 300 m provides a very rich dataset for analysis.
“For all the missions, one of the fundamental things we tried to have was maximum coverage of the total station for the duration of the mission,” Tuna said. “Because the total station is a fundamental backbone of our ground-truth, our ground-truth position measurements came directly from the total station through the AP20. Whether or not we had GNSS along the route, we wanted to ensure that at least at the beginning and the end of every mission, we had measurements from the total station.”
Paths Forward
“When we were initiating GrandTour, we were really adamant that it must be a long-term project,” Tuna said. “We knew this because we had some understanding of how hard it is to do a successful mechatronics project at this scale that works, but it was harder than even we expected. We are where we are now after two years of painstaking development and fine-tuning to ensure maximum accuracy. Now, we must ask ourselves what exact data is needed next data that is even more purposeful than before and aimed at solving more of the problems researchers are facing today.”
GrandTour is not a one-off. “The project continues, and it will become even more sophisticated,” Tuna said. “We are at the stage of figuring out the next iteration of the project and, most importantly, what other kinds of platforms we can put Boxi on. This could include other quadruped robots beyond ANYmal, wheeled systems, tracked vehicles and more. We are hardware- and platform-agnostic.”
The uptake from the R&D community has been rapid and growing. “Today, we have at least 330 live users for this dataset,” said Tuna. “These are only the users we can see, those who have requested access, but as the data is open source, so there will be many more.”
I asked about other entities: other groups might want to do their own training-data runs to add to this open dataset. Would they have to build their own Boxi to the same specifications as the GrandTour team, or could they borrow one? “I can answer this in two ways,” responded Tuna. “We selected the sensors and modalities for Boxi from off-the-shelf solutions as much as possible. For example, if you can get the same camera, you would basically have the same image type that you use to train with the GrandTour project assets. On the other hand, we also released the Boxi hardware design, as well as the electronics and electrical sets, so if people want to, they can recreate it.”
Tuna further explained that the main purpose of Boxi was to capture data from a wide range of sensors, often with redundancy, from which researchers could analyze different modalities and combinations. However, they may not need as many sensors, or they may need to add others; the training data nonetheless provides a rich foundation. “Some researchers and developers look at this array of sensors and data to help determine which are the right ones for their needs. You do not always know which one is right for you, so you can select the right bundle and then, in your own institute, decide which subset of sensors you would actually use,” he said.
“Hence, with the Boxi research paper, we put quite a lot of emphasis on this,” Tuna added. “One of the main purposes of GrandTour was to showcase what kinds of modalities your solution might need, and then, from there, how to create your own solution. This may only need to be a subset of what is included on Boxi, perhaps something more streamlined, more compact, or cheaper.”
There was always the intention of eventually doing longer missions, and missions that could somehow encompass what surveyors might call “side loops.” This would be explorations of areas branching off from what would be decision points that would occur in different use case scenarios. There’s also the need to keep up on the types of sensors and approaches constantly being developed and updated for reality capture, robotics and autonomy.
“For GrandTour and Boxi, we are always considering updates to our sensor suite,” said Tuna. “We also want it to encompass enough capability to support niche targets that might otherwise matter only to certain communities.” One example is recent advances and products in the field of RGB (native color) lidar, from firms like Ouster. Small-format, full-data radar is another area of interest. And just imagine when things such as small-format quantum lidar/radar/cameras come to fruition.
What does the future hold? “This solution became more than just a side project that dies after initial data collection,” Tuna said. “It was developed in such a way that it works with essentially just a click. That is the kind of convenience that will really benefit researchers and makes it a tremendous resource for developing navigation and autonomy solutions. So, when we told colleagues that this solution emits data from all the sensors and can be used live, they began doing more camera–lidar fusion, inertial odometry, and mapping.”
Beep on!