Skip to content
Main Site News Console

Overhyped Visuotactile Technology: An Easily Replicated Business—What’s Its Investment Value?

· 量子位
国内AI

Overhyped Vision-Based Tactile Sensing: An Easily Replicated Business—How Much Investment Value Does It Really Have?

Written by Yunzhong, from Aofeisi Temple

The embodied intelligence industry is entering a critical phase of value realization. Expectations for embodied intelligence have advanced toward deployment and value creation in real-world industrial scenarios centered on physical interaction.

The core demand for physical interaction makes tactile sensing the next-generation core sensing technology after vision.

This sensing system uses force sensors as its core sensing units, expanding outward into high-precision, highly integrated end-of-arm tactile sensors, and further into curved, conformal, large-area distributed electronic skin. These technologies together form a complete closed loop, taking robots from fine force control at the fingertips to full-body physical interaction.

The industry’s focus is also shifting away from competition based on isolated performance metrics such as precision and accuracy, toward an industrial infrastructure narrative centered on long-cycle, all-scenario operational stability.

Among the mainstream tactile-sensing technologies currently available, vision-based tactile sensors (VBTS), which rely on photoelectric conversion, are undoubtedly one of the most popular directions. The field is crowded, and it continues to attract substantial attention from both academia and industry.

Yet beneath all the hype, is this technical approach really capable of shouldering the role of industrial-grade infrastructure required for the large-scale deployment of embodied intelligence?

Let us objectively examine its technical reality from the perspectives of technological origins, underlying principles, performance limits, and deployment feasibility.

This is the first article in a series analyzing tactile-sensing technology routes. Centered on the requirements for the large-scale deployment of embodied intelligence, the series will systematically examine current mainstream and emerging tactile-sensing approaches, including vision-based tactile, piezoresistive, capacitive, and magnetoelectric technologies. It will compare them across dimensions including technical principles, engineering reliability, integration forms, and industrial deployment capabilities.

The goal is to clarify the critical stage at which embodied intelligence is entering real-world industrial scenarios, as well as the true capability boundaries and applicable scenarios of different technical routes.

The Divide Between Academic Prosperity and Industrial Feasibility

1. A Low-Barrier Frenzy Fueled by the “Vision Dividend”

Since the GelSight era, vision-based tactile sensors have seen some optimization in thickness and durability, but their underlying architecture still relies on imaging with compact CMOS camera modules. Their overall capabilities have not undergone any fundamental improvement.

Because the vision-sensor industry is already highly mature, CMOS chip development has progressed relatively slowly over the past decade and a half, particularly in the video field. Vision-based tactile products largely reuse mature solutions similar to those used in video cameras. Such components are readily available from Taobao or solution providers, leaving their core specifications—such as spatial resolution and sampling frequency—largely unimpressive, with no fundamental breakthroughs.

△ Visualization by the GelSight team in 2009

2. Academia: A Research Boom Driven by the Spillover of Vision Algorithms

After 2010, as vision algorithms achieved explosive breakthroughs at top conferences such as CVPR, researchers at robotics conferences including ICRA and IROS were able to directly repurpose large numbers of existing vision algorithms for processing vision-based tactile data.

Take the three most commonly used algorithms as examples:

  • ResNet/CNN—general-purpose feature-extraction backbones from the vision field. Vision-based tactile research often connects them directly to tactile images and fine-tunes them for force classification or material identification. The reuse pathway is relatively straightforward.

  • U-Net—originally designed for medical image segmentation. In vision-based tactile sensing, it is widely used to predict three-dimensional shapes or force maps from deformation maps of elastomers, with relatively little effort devoted to architecture improvements tailored to the physical characteristics of touch.

  • Optical Flow—a classic computer-vision technique for calculating pixel motion. It has been transferred to tracking marker displacement and indirectly estimating applied forces, making it one of the few approaches with some connection to tactile physics. However, its core algorithm still comes entirely from the vision field, and it depends on the presence of markers, limiting its range of applicability.

These three algorithms cover a large number of common research topics in vision-based tactile sensing. Overall, the field relies heavily on mature tools from the vision ecosystem at the algorithmic level. It has yet to establish an original algorithmic paradigm independent of vision and grounded in the physical essence of tactile perception. This “borrowed-is-best” approach may accelerate paper publication, but it also means that vision-based tactile sensing has made almost no groundbreaking algorithmic contributions independent of the vision ecosystem.

3. Industry: Low Barriers, High Homogeneity, and Investment Risk

Vision-based tactile-sensor developers have achieved no substantive breakthroughs in either underlying hardware or core algorithms. In essence, they do not control the full stack of proprietary technology, and their supply chains are also highly dependent on external suppliers, creating a significant risk of lacking independent control.

The core component—the CMOS image sensor—is dominated by a small number of international giants. Fewer than 10 companies in China can independently supply CMOS sensors. The source of this core-component supply chain is not in the hands of these vision-based tactile startups. What is called “R&D” is largely system integration built on other companies’ finished products. Vision-based tactile manufacturers mostly adopt off-the-shelf camera-module solutions and lack control over the underlying manufacturing processes of key sensors.

Moreover, the physical construction of a vision-based tactile sensor is extremely simple: a miniature camera, a light source, a transparent elastomer, and markers are sufficient to build a prototype.

This “take-and-use” approach to the mature vision supply chain makes the hardware-development barrier extremely low. University students in China can develop prototypes independently in a short period of time and sell them cheaply on e-commerce platforms, fully exposing the hollowed-out nature of the hardware technology in this field.

△ Camera imaging system

△ Vision-based tactile sensors directly use customized camera-module solutions; the left image shows a SONY IMX-series CMOS camera module

Low algorithmic barriers and highly reusable hardware have directly lowered the threshold for starting a business, encouraging large numbers of followers to enter the field. Vision-based tactile startups have emerged in large numbers both in China and abroad, making this the tactile-sensing segment with the most entrants.

Dozens of companies are developing products based on the same technical principles and the same sources of open-source algorithms. As a result, the sensors they deliver have difficulty establishing meaningful differences in core performance metrics.

The vision-based tactile sector has long exhibited a competitive landscape of “numerous vendors and intense homogenized competition.” To date, no leading company with representative technology and market dominance has emerged. From an investor’s perspective, a sector whose core technology is not independently controlled, whose moat is unclear, and whose business model can be easily replicated inherently faces challenges in delivering predictable long-term investment returns.

A company’s ultimate moat lies in the pricing power generated by core technology—not in relying on economies of scale to earn hard-won profits from low-value-added activities. For this kind of assembled innovation, characterized by “outsourced supply chains and purchased algorithms,” capital will view it only as a high-risk, low-barrier asset. It lacks the ability to define industry standards, let alone build an insurmountable commercial barrier.

Therefore, smart capital will place its bets only on those genuine disruptors capable of breaking through physical limits, controlling core technologies across the full stack, and tackling truly difficult problems.

In other words, the popularity of vision-based tactile sensing is fundamentally built on the spillover of vision algorithms and low hardware barriers—not on a fundamental breakthrough in tactile measurement. As a derivative technical route, it lacks an independent industrial foundation for becoming the infrastructure of robotic touch from the very beginning.

Industrial Deployability: Hardware Architecture and Long-Term Reliability

Whether a tactile-sensing solution can support large-scale deployment depends on two unavoidable hard barriers: hardware reliability and service life.

1. Minimum Focal-Length Constraints and Structural Thickness

To ensure the accuracy of deformation reconstruction and image uniformity, the camera and the elastomer’s contact surface must meet minimum working-distance requirements. Space must also be reserved for the illumination needed to achieve uniform lighting.

In current mainstream solutions, sensors using lens designs are generally more than 10 mm thick. The overall systems are bulky and offer a low degree of modularity. Even lensless designs used to reduce thickness introduce optical-distortion issues and still cannot escape their dependence on CMOS photosensitive elements, light sources, and elastomers.

For dexterous hands evolving toward high degrees of freedom and miniaturization, the interior of each fingertip is already densely packed with essential components such as joint motors, reducers, and tendon drives. Available space is extremely limited. The rigid volume requirements of vision-based tactile modules further compress the mechanical system’s design space, forcing developers to compromise among fingertip degrees of freedom, load capacity, and tactile-sensing accuracy.

△ Source: Tactile Image Sensors Employing Camera: A Review

2. Rigid Form Factor and the Deployment Limits of Electronic Skin

Vision-based tactile sensors rely on enclosed optical cavities and rigid imaging substrates to maintain stable imaging. They therefore do not naturally support flexible extension or conformal attachment to curved surfaces.

This imposes clear limits on their deployment form: they can only be deployed as independent rigid modules on flat or low-curvature contact areas, such as robotic grippers and dexterous fingertips. They cannot flexibly conform to complex body-surface structures such as a robot’s torso, joints, or curved outer shell in the manner of electronic skin.

This limitation fundamentally caps the scale of deployment. The technology can serve only as a localized end-effector sensing solution and cannot support the construction of a full-body tactile-sensing system for robots.

3. Integration and Cost Cascade Effects of Multi-Camera Architectures

When tactile sensing expands from a single fingertip to multiple fingers on both hands, the underlying deficiencies of vision-based tactile sensing escalate from technical bottlenecks into a systemic breakdown.

Using a typical configuration of two hands, five fingers, nine finger pads, and two palms as an example, a single dexterous hand would need nearly 30 vision-based tactile modules—equivalent to forcibly cramming nearly 30 camera modules into mechanical structures where every millimeter is precious.

Synchronized acquisition of multiple video streams and complex cable routing directly conflict with the precision transmission mechanisms inside a dexterous hand, making integration extremely difficult.

△ The image shows the architecture of a monitoring system for parallel transmission of multiple HD video streams, which requires substantial physical-resource redundancy

More devastatingly, the increase in visual signal channels directly triggers a “death spiral” in computing requirements and costs: 30 concurrent channels require more than 160 TOPS of computing power, necessitating an external server, while power consumption exceeds 100 W.

This also completely breaks the latency limit for millisecond-level real-time control. Bandwidth pressure forces the system to significantly reduce image quality in order to maintain real-time performance. Force inversion must pass through the entire pipeline of “acquisition → decoding → inference → mapping,” resulting in end-to-end latency of tens of milliseconds, while the closed loop for compliant control requires only 1–4 ms. This directly consumes the high-definition imagery on which vision-based tactile sensing depends.

This chain reaction of “more visual channels → greater algorithmic complexity → increased computing requirements → higher costs → restricted deployment” exposes a fundamental architectural deficiency of vision-based tactile sensing: it transforms what should be a lightweight, low-power tactile endpoint into a video-monitoring system that must be sustained by a GPU cluster.

On a robotic platform that is extremely sensitive to size, power consumption, real-time performance, and cost, this is a complete contradiction.

Finally, more than 30 cameras mean more than 30 independent power and data cables. The interior of a dexterous hand is already occupied by motors and reducers, while the cables must be routed through hollow modules at the base of the fingers with diameters of only a few millimeters—there is simply no room for 30 cable bundles to pass through simultaneously.

Even if the cables were forcibly routed, the dense bundles would rub against and interfere with the precision transmission mechanisms, drastically reducing mechanical service life and causing signal crosstalk. The visual pathway reaches a dead end at the very first step: cable routing.

The result is an endless loop: the cable bundles cannot be routed out, the computing load cannot be sustained, power consumption cannot be supported, costs rise, and deployment becomes restricted. More visual channels → no viable cable routing → increased computing requirements → exploding power consumption → higher costs → restricted deployment. Vision-based tactile sensing turns a lightweight tactile endpoint into a GPU-dependent video-monitoring system, and it does not work on the robot itself.

4. The “Zero-Sum Game” of Elastomer Materials

Vision-based tactile sensors face an unavoidable fundamental physical dilemma: the softness or hardness of the elastomer directly determines both the system’s sensing capabilities and its service life, and the two cannot be achieved simultaneously.

△ Gel materials used in the GelSight vision-based tactile sensor

The sensing capability of a vision-based tactile sensor is directly causally related to the softness or hardness of its surface elastomer: the softer the elastomer, the higher the system’s sensitivity and the clearer the imaging; conversely, the harder the elastomer, the lower the sensitivity and the poorer the image quality.

Specifically, a soft elastomer undergoes more significant deformation when it contacts an object, amplifying tiny differences in surface texture into visible displacements that the camera can detect.

Therefore, the softer the material, the more sensitive it is to microscopic deformation, the clearer the imaging, and the richer the texture details.

However, once the elastomer is hardened to improve durability or force-measurement linearity, or once the reflective layer is thickened, the magnitude of deformation shrinks dramatically. Microscopic textures can barely produce sufficient displacement on a hard surface, causing the contrast of the images captured by the camera to decrease and details to blur. The sensor’s sensitivity to subtle contact also declines sharply.

It must be understood that the essence of vision-based tactile sensing is a physical “zero-sum game”: a soft elastomer is necessary for obtaining clear images, but it is also the source of rapid system aging.

Once the material is hardened for industrial durability, microscopic imaging capability effectively drops to zero. This deadlock—“soft means sensitive, hard means dull”—means that in real production lines, the sensor will inevitably face continuous drift in its optical characteristics due to creep and wear. Long-term reliability is therefore out of the question.

When a technology cannot be miniaturized and integrated into a complete dexterous hand—including the finger pads and palms—cannot cover the whole body, and cannot operate stably in harsh environments, the industrial expectations placed upon it have clearly been severely overdrawn.

The Underlying Logic of Dexterous Manipulation: Force, Not Texture

1. Two-Dimensional Projection and an Underdetermined Inverse Problem

The core mechanism of vision-based tactile sensing is to compress the three-dimensional deformation of the outermost elastomer into a two-dimensional pixel image through an embedded camera, and then rely on algorithms to reconstruct the contact-force distribution.

From the perspective of physical measurement, this approach does not directly measure force based on the material’s intrinsic effects. Instead, it couples and compresses normal indentation and tangential shear deformation in three-dimensional space into the features of a single two-dimensional image.

This is essentially a typical underdetermined inverse problem: the system must forcibly reconstruct three-dimensional force components from a two-dimensional observation with reduced information.

Under complex conditions such as multipoint and edge contact, error accumulation in photometric stereo, together with the physical entanglement of normal- and tangential-force information, leads to a significant increase in force-decoupling errors. Such errors arise because the original two-dimensional displacement field lacks information relative to the three-dimensional surface contact-force field, making them difficult to fundamentally eliminate through algorithmic optimization.

2. Image Resolution Does Not Equal Force-Measurement Resolution

A common misconception in the industry is to equate the number of optical pixels with tactile spatial resolution.

In reality, pixels are merely sampling units for optical imaging, not the number of independent force-measurement points.

The effective spatial resolution of tactile sensing is subject to hard upper limits imposed by three physical constraints: the attenuation characteristics of deformation transmission within the elastomer, the non-uniqueness of algorithmic inversion, and the limits of force-decoupling capability described above.

The actual density of effective force-measurement points is far lower than the pixel density of the CMOS sensor, as shown below. Simply piling on more pixels does not equivalently improve tactile perception; instead, it can degrade the signal-to-noise ratio (SNR) by introducing redundant optical noise.

△ The number of pixels does not equal the number of actual physical measurement points

The effective spatial resolution of vision-based tactile sensing is also locked by the hard upper limits imposed by three physical constraints:

  • Viscoelastic effects (the viscoelasticity of transparent gels causes hysteresis, making the mapping between displacement and force time-dependent);

  • Non-uniqueness of algorithmic inversion (the same image may correspond to multiple force states; algorithms can only “guess,” not “solve”);

  • Limits of force-decoupling capability (normal force, shear force, torque, and regularization terms interfere with one another and cannot be perfectly separated from a two-dimensional image).

3. Structural Contradiction: Sensitivity and Industrial Durability Cannot Coexist

However, vision-based tactile sensing faces a structural contradiction: to capture fine textures, the transparent gel (elastomer) must be sufficiently soft to amplify deformation signals, but this directly sacrifices durability; for long-term use, both the surface and interior of the transparent gel must be sufficiently durable to ensure basic resistance to damage and light leakage.

The vision-based tactile route contains an irreconcilable “inherent deadlock” at the level of its fundamental physical architecture: the elastomer’s mechanical properties mean that the solution can choose only between “highly sensitive” and “able to survive”; there is no third option.

If the elastomer is hardened to improve reliability in industrial environments, fine texture information will be lost or substantially attenuated. Yet texture is one of the few sources of differentiated value offered by vision-based tactile sensing.

This also exposes the inherent contradiction in the underlying physical architecture of the vision-based tactile route: the system cannot capture microscopic textures with high sensitivity while also ensuring high industrial reliability.

To capture microscopic textures and fine deformation, the transparent gel elastomer must be extremely soft. But this amounts to condemning the sensor to the fate of a “consumable.” In industrial environments characterized by oil contamination, high temperatures, and mechanical wear, soft gels are not only highly susceptible to aging and tearing; they also experience signal drift due to hysteresis.

Conversely, if the elastomer is hardened for industrial-grade reliability, the deformation signals of microscopic textures are forcibly flattened. The “high-resolution tactile perception” that vision-based tactile sensing prides itself on instantly becomes a low-end product no different from an ordinary force sensor.

This is not merely a challenge in materials science; it is also the bankruptcy of the commercial logic: vision-based tactile sensing attempts to use a “delicate laboratory material” to solve a “rough industrial-site problem.”

This fundamental contradiction in the underlying physical architecture means that its product forms will forever oscillate between “precise but fragile” and “rugged but mediocre.” It cannot produce a truly industrial-grade standard product that combines high sensitivity with long service life.

Therefore, vision-based tactile sensing struggles to provide deterministic force feedback during dexterous manipulation and cannot simultaneously meet the high reliability and consistency required for industrial deployment.

The durability and health of a vision-based tactile sensor must never be defined solely by whether its camera continues to operate.

The system’s “survival” merely indicates that the optical acquisition path remains intact. Its true sensing lifetime depends on the physical degradation of the elastomer. Once the elastomer undergoes irreversible deformation due to wear, creep, or aging—causing changes in consistency and accuracy—its force-to-optical mapping has already become distorted. At that point, even if the camera is still capturing images normally, its output is merely a stream of “error information.”

For robots that rely on closed-loop feedback, an inaccurate force signal is more damaging than no signal at all. It can cause the control strategy to make incorrect judgments, resulting in manipulation failures or even hardware damage.

Therefore, measurement inaccuracy caused by elastomer failure is the only valid standard for determining poor sensor durability. Simply keeping the camera powered on is merely false survival under a “visual illusion.”

Force Accuracy: The Gap in Dynamic Performance During Engineering Deployment

1. Insufficient Force-Sensing Sensitivity

The core mechanism of the vision-based tactile approach is to use a camera to capture the optical deformation of an elastomer under force, and map the displacement of marker points or random speckles to a force distribution.

In practice, force estimation depends on optical-deformation modeling and image-processing algorithms. The mapping relationship is complex and easily affected by environmental interference.

Force inversion depends entirely on optical-deformation modeling and image processing. Yet the elastomer’s nonlinear hysteresis, illumination drift, and material aging continuously distort this mapping relationship.

More critically: the displacement of marker points caused by changes in pressure is concentrated primarily in the edge pixels of the contact area, while the central region barely moves. Under the rated load, displacement at the edge of the contact area is negligible—the information that the algorithm can use to “measure-incomplete output.