機械人進步主要來自 AI,架構複用語言模型積木The progress of robots mainly comes from AI, using architecture to reuse language model building blocks.
Construction Physics 拆解驅動機械人嘅視覺語言動作模型,以一個開放權重模型為例:圖像、文字指令同機械人狀態一次過編碼,再經語言骨幹加動作模組反覆去噪,最後輸出五十個動作。視覺編碼器四億參數、語言模型二十億參數、動作模組三億參數,架構同聊天模型高度重疊。Construction Physics breaks down the visual-language action model driving robots. Taking an open-weight model as an example: images, text instructions, and robot states are encoded all at once, then repeatedly denoised through a language backbone and action module, finally outputting fifty actions. The visual encoder has 400 million parameters, the language model 2 billion parameters, and the action module 300 million parameters, with an architecture highly overlapping with chat models.
It’s been over a year since I last looked at the state of humanoid robots, and enthusiasm for the technology shows no sign of waning. Startups NEURA Robotics and Figure AI raised $1.4 billion and $1 billion, respectively, and Apptronik raised $520 million. China’s Unitree went public, raising roughly $900 million, and Agility Robotics plans to go public via SPAC later this year. A few weeks ago China held the second annual World Humanoid Robot Games, showing humanoid robots performing all manner of impressive physical feats.
We’ve also seen some impressive-looking robotics demos over the last year, though these must be evaluated with a large grain of salt. Figure showed off its robots sorting packages for hours at a time and autonomously unloading a dishwasher. Physical Intelligence showed off its robot model making coffee and folding boxes in a chocolate factory. Generalist AI showed off its GEN-1.5 model learning new tasks from just a few examples.
Some progress in robots is a function of hardware advances: better actuators, and so on. But the lion’s share of it is due to advances in robot AI, using specially developed AI models to control a robot. As more general AI continues to rapidly advance in capabilities, one obvious question is whether we’re going to see the same sort of acceleration in robotic AI. Right now robotic AI capabilities still seem very limited, but given how fast more general AI improved, I can imagine this changing very quickly. Because of this, it’s worth understanding how the AI models being developed to control robots work.
There are a few different sorts of robotic AIs, which are sometimes referred to as “policies” (where a “policy” is something that maps a particular set of inputs — robot state, sensor data, instructions it’s been given — to a set of robot actions). For this essay, we’ll look at one commonly used architecture, employed by companies such as Figure, Unitree, Physical Intelligence, and Nvidia: the vision-language-action model (VLA). More specifically, we’ll look at one popular VLA, Physical Intelligence’s open-weight π0.5 VLA, which was released in 2025 and has become widely used (though it’s not Physical Intelligence’s most advanced model).
While a large language model takes text as an input and gives text as an output, a VLA takes text, images, and robot information as an input and gives a series of robot actions as an output. It does this using many of the same components that made LLMs so successful, specifically attention and the transformer architecture. It’s not amazingly obvious if VLA’s will remain the primary robotic AI paradigm, but as of now they’re a very common method for controlling robots.
Linear algebra basics
To learn how AI works (be it a VLA or anything else), it’s useful to know a very small amount of linear algebra. Specifically, we want to understand a few different operations for manipulating arrays of numbers, since this is most of what AI models do.
For instance, say we have the following two-dimensional arrays (or matrices) of numbers:
One thing we can do to these arrays is add them together. This works exactly like how you’d expect: each number in one array is added to the corresponding number in the other array, giving you a new array with all the resulting additions.
Another thing we can do is multiply or divide an array by a single value. This also works more or less like you’d expect: each value in the array gets multiplied or divided by the respective value.
But what if we want to multiply two matrices together? To do this, we need an operation called the dot product. The dot product takes two one-dimensional arrays of numbers (also known as vectors), multiplies each value in one vector by the corresponding value in the other vector, then sums the result.
You can think of the dot product as measuring two things: how large two vectors are, and how similar they are to each other. (If you normalize the vectors, scaling the values so everything is between 0 and 1, then the dot product is entirely a measure of how similar two vectors are to each other.)
When we multiply two matrices together, we’re simply doing a bunch of dot products: each value in the resulting matrix is the dot product of the row of one array and the column of another. Because the dot product requires two lists of equal length, multiplication of matrices must be done in a certain way: the number of columns of one matrix must be equal to the number of rows of the other matrix. And the size of the output matrix will be a function of the sizes of the input matrices.
I find it easiest to understand matrix multiplication by putting one matrix on the left and the other matrix on the upper right. The output matrix fits in the space between them, each value the dot product of the first matrix’s rows and the second’s columns:
Neural network basics
Most modern AI models, as we know, are built using neural networks. The classic image of a neural network is something like this:
You have some series of input neurons that represent your input data, which are connected to various intermediate layers of neurons, which then get connected to output neurons. Depending on what data is fed into them and how the connections between different neurons have been set, neurons will “activate” to various degrees, taking some input from the neurons in the previous layer and sending some output along to the neurons in the next layer. This structure is sometimes called a “multilayer perceptron,” or MLP.
However, I think it’s much easier to understand neural networks by looking at how they’re actually implemented, which is done by multiplying arrays of numbers (in machine learning, these arrays are sometimes called “tensors,” where a tensor is just an array of numbers that can be any dimension. Google, for instance, has a machine-learning library called “TensorFlow”). The neural network above, for instance, would actually be implemented via something like this:
The three input neurons become a one-dimensional array, or vector, with three numbers in it, which is our input data. This vector then gets passed to the hidden layer, where it goes through several steps. First, it gets multiplied by a two-dimensional matrix (W1 in the figure). This matrix has four rows (corresponding to the four neurons in the original diagram) and three columns (one for each value in the input vector): it takes a vector three numbers long (or 3-vector) as an input and spits out one four numbers long as an output. The resulting vector of four numbers is then added to another 4-vector (b1 in the figure), the values of which are known as “biases.” The weights and biases of the various layers are known as the “parameters” of the network and are what get modified when the neural network is being trained.
Once the biases have been added, the resulting vector (known as a “pre-activation,” z in the figure) then goes through what’s called a nonlinearity, which is just a function that does different transformations to an input depending on what that input is. A common nonlinearity, which the example above uses, is “ReLU,” for rectified linear unit. All ReLU does is replace any negative numbers in the vector it’s given with zeroes. Other nonlinearities are the sigmoid (an S-shaped function that squeezes values to be between 0 and 1) and GELU (a somewhat ReLU-like function that’s smooth instead of sharply kinked).
In a basic MLP like this one, each hidden layer in a neural network will consist of this matrix multiplication, then bias, then nonlinearity. The network above has just one hidden layer, so after the nonlinearity the resulting vector is then passed to the output neurons. These transform the data again using another matrix multiplication (W2) and bias addition (b2), yielding a vector of two numbers as the output (y).
……(原文過長,此處截斷)
自从我上次关注人形机器人已经一年多了,对这项技术的热情仍然没有减退。初创公司 NEURA Robotics 和 Figure AI 分别筹集了 14 亿美元和 10 亿美元,而 Apptronik 筹集了 5.2 亿美元。中国的 Unitree 上市,筹集了大约 9 亿美元,Agility Robotics 计划今年晚些时候通过特殊目的收购公司(SPAC)上市。几周前,中国举办了第二届世界人形机器人运动会,展示了人形机器人完成各种令人印象深刻的体能表演。
在过去的一年里,我们还看到了一些看起来很厉害的机器人演示,尽管这些必须抱有很大的怀疑态度来评估。Figure 展示了其机器人连续数小时分拣包裹,并自主清空洗碗机。Physical Intelligence 展示了其机器人模型在巧克力工厂中制作咖啡和折叠盒子。Generalist AI 展示了其 GEN-1.5 模型仅通过几个示例就学习新任务。
机器人方面的一些进展是硬件进步的结果:更好的执行器等等。但绝大部分进展归功于机器人 AI 的发展,使用专门开发的 AI 模型来控制机器人。随着更通用的 AI 能力持续快速提升,一个显而易见的问题是,我们是否会看到机器人 AI 出现同样类型的加速。目前机器人 AI 的能力仍显得非常有限,但考虑到更通用的 AI 提升得如此之快,我可以想象这种情况会很快改变。因此,了解用于控制机器人的 AI 模型的工作原理是值得的。
有几种不同类型的机器人人工智能,有时被称为“策略”(“策略”是指将特定输入集合——机器人状态、传感器数据、所给指令——映射到一系列机器人动作的东西)。在本文中,我们将研究一种常用的架构,这种架构被 Figure、Unitree、Physical Intelligence 和 Nvidia 等公司采用:视觉-语言-动作模型(VLA)。更具体地说,我们将研究一种流行的 VLA,即 Physical Intelligence 的开放权重 π0.5 VLA,它于2025年发布,并已被广泛使用(尽管它不是 Physical Intelligence 最先进的模型)。
While a large language model takes text as input and gives text as output, a VLA takes text, images, and robot information as input and gives a series of robot actions as output. It does this using many of the same components that made LLMs so successful, specifically attention and the transformer architecture. It is not entirely clear whether VLAs will remain the primary paradigm for robotic AI, but for now, they are a very common method for controlling robots.
Linear algebra basics
要了解人工智能是如何工作的(无论是 VLA 还是其他什么),了解少量的线性代数是有用的。具体来说,我们希望理解一些操作数字数组的不同方法,因为这几乎涵盖了人工智能模型的大部分工作。
For instance, say we have the following two-dimensional arrays (or matrices) of numbers:
One thing we can do with these arrays is add them together. This works exactly as you would expect: each number in one array is added to the corresponding number in the other array, giving you a new array with all the resulting sums.
Another thing we can do is multiply or divide an array by a single value. This also works more or less like you’d expect: each value in the array gets multiplied or divided by the respective value.
But what if we want to multiply two matrices together? To do this, we need an operation called the dot product. The dot product takes two one-dimensional arrays of numbers (also known as vectors), multiplies each value in one vector by the corresponding value in the other vector, then sums the result.
You can think of the dot product as measuring two things: the magnitude of two vectors and how similar they are to each other. (If you normalize the vectors, scaling the values so that everything is between 0 and 1, then the dot product becomes entirely a measure of how similar the two vectors are to each other.)
当我们将两个矩阵相乘时,我们实际上只是在做一系列点积:结果矩阵中的每个值都是一个矩阵的行和另一个矩阵的列的点积。由于点积需要两个长度相等的列表,因此矩阵的乘法必须以特定的方式进行:一个矩阵的列数必须等于另一个矩阵的行数。输出矩阵的大小将取决于输入矩阵的大小。
I find it easiest to understand matrix multiplication by putting one matrix on the left and the other matrix on the upper right. The output matrix fits in the space between them, each value being the dot product of the first matrix’s rows and the second matrix’s columns:
Neural network basics
正如我们所知,大多数现代人工智能模型都是基于神经网络构建的。神经网络的经典形象是这样的:
你有一些表示输入数据的输入神经元,它们连接到各种中间层的神经元,然后再连接到输出神经元。根据输入到它们的数据以及不同神经元之间的连接方式,神经元会以不同程度“激活”,从前一层的神经元接收一些输入,并将一些输出发送到下一层的神经元。这种结构有时被称为“多层感知器”,或 MLP。
However, I think it’s much easier to understand neural networks by looking at how they’re actually implemented, which is done by multiplying arrays of numbers (in machine learning, these arrays are sometimes called “tensors,” where a tensor is just an array of numbers that can be any dimension. Google, for instance, has a machine-learning library called “TensorFlow”). The neural network above, for instance, would actually be implemented via something like this:
三个输入神经元变成一个一维数组,或者向量,其中包含三个数字,这就是我们的输入数据。然后,这个向量会传递到隐藏层,在那里经历几个步骤。首先,它会与一个二维矩阵(图中的 W1)相乘。这个矩阵有四行(对应原始图中的四个神经元)和三列(输入向量中的每个值对应一列):它以一个长度为三的向量(或3-向量)作为输入,并输出一个长度为四的向量。得到的四个数字的向量随后会与另一个四维向量(图中的 b1)相加,其数值被称为“偏置”。各层的权重和偏置被称为网络的“参数”,也是在训练神经网络时会被修改的内容。
一旦偏置被加上,得到的向量(在图中称为“激活前值”,用 z 表示)随后会经过一种被称为非线性函数的处理,这只是一种根据输入不同对其进行不同变换的函数。一个常见的非线性函数(上例中使用的)是“ReLU”,即修正线性单元。ReLU 的作用就是将向量中所有负数替换为零。其他的非线性函数有 sigmoid(一个 S 形函数,将值压缩到 0 到 1 之间)和 GELU(类似 ReLU 的函数,但比起锐角更平滑)。
In a basic MLP like this one, each hidden layer in a neural network consists of a matrix multiplication, followed by a bias addition, and then a nonlinearity. The network above has only one hidden layer, so after the nonlinearity, the resulting vector is passed to the output neurons. These output neurons transform the data again using another matrix multiplication (W2) and bias addition (b2), producing a vector of two numbers as the output (y).
……(The original text is too long, truncated here)
原文出處:Source: Construction Physics ↗