上一篇【第33篇】常量指令——把数字推到栈上的艺术
下一篇【第35篇】栈指令——dup/swap 的奇妙世界
摘要
上一篇我们造了常量指令——把值无中生有地推到操作数栈上。但 JVM 里的数据还有一个更重要的来源:局部变量。
这一篇要造的是加载指令(loads)和存储指令(stores),它们是局部变量表和操作数栈之间的唯一桥梁:
局部变量表 操作数栈 LocalVars xload OperandStack ┌────────┐ ────────► ┌────────┐ │ slot 0 │ │ │ │ slot 1 │ ◄──────── │ 栈顶 │ │ slot 2 │ xstore │ │ └────────┘ └────────┘JVM 规范给这两类各分配了33 条指令,合计 66 条——占全部 205 条指令的32%。本章实现其中的 50 条(数组相关的xaload/xastore留到第 8 章)。
读完这一篇,你会理解:为什么局部变量索引只有 0~3 有快捷指令?为什么iload_3存在而iload_4不存在?超过 255 个局部变量怎么办?
一、加载指令:局部变量表 → 操作数栈
1.1 33 条加载指令的全景
加载指令从局部变量表获取变量,然后推入操作数栈顶。按照所操作变量的类型可以分为6 类,opcode 从0x15到0x35,恰好连续 33 个:
opcode 助记符 说明 本章 ────────────────────────────────────────────────── 0x15 iload int,索引由操作数给出 ✅ 0x16 lload long ✅ 0x17 fload float ✅ 0x18 dload double ✅ 0x19 aload reference ✅ ────────────────────────────────────────────────── 0x1A iload_0 int,索引隐含在操作码 ✅ 0x1B iload_1 ✅ 0x1C iload_2 ✅ 0x1D iload_3 ✅ 0x1E lload_0 long ✅ 0x1F lload_1 ✅ 0x20 lload_2 ✅ 0x21 lload_3 ✅ 0x22 fload_0 float ✅ 0x23 fload_1 ✅ 0x24 fload_2 ✅ 0x25 fload_3 ✅ 0x26 dload_0 double ✅ 0x27 dload_1 ✅ 0x28 dload_2 ✅ 0x29 dload_3 ✅ 0x2A aload_0 reference ✅ 0x2B aload_1 ✅ 0x2C aload_2 ✅ 0x2D aload_3 ✅ ────────────────────────────────────────────────── 0x2E iaload 数组元素加载(int) ❌ 第 8 章 0x2F laload ❌ 0x30 faload ❌ 0x31 daload ❌ 0x32 aaload ❌ 0x33 baload ❌ 0x34 caload ❌ 0x35 saload ❌规律极其整齐:
| 类型 | 带索引版 | 快捷版 | 数组版 |
|---|---|---|---|
| int | iload | iload_0~iload_3 | iaload |
| long | lload | lload_0~lload_3 | laload |
| float | fload | fload_0~fload_3 | faload |
| double | dload | dload_0~dload_3 | daload |
| reference | aload | aload_0~aload_3 | aaload |
| byte/char/short | — | — | baload/caload/saload |
5 × (1 + 4) + 8 =33✅
1.2 代码实现:以 iload 系列为例
在ch05/instructions/loads目录下创建iload.go文件:
packageloadsimport"jvmgo/ch05/instructions/base"import"jvmgo/ch05/rtda"// Load int from local variabletypeILOADstruct{base.Index8Instruction}typeILOAD_0struct{base.NoOperandsInstruction}typeILOAD_1struct{base.NoOperandsInstruction}typeILOAD_2struct{base.NoOperandsInstruction}typeILOAD_3struct{base.NoOperandsInstruction}注意这里两种抽象指令的混合使用:
ILOAD嵌入base.Index8Instruction——索引来自操作数(1 字节),需要FetchOperands去读ILOAD_0~ILOAD_3嵌入base.NoOperandsInstruction——索引写死在类型里,没有操作数
1.3 用辅助函数消除重复
为了避免重复代码,定义一个函数供 5 条指令共用:
func_iload(frame*rtda.Frame,indexuint){val:=frame.LocalVars().GetInt(index)frame.OperandStack().PushInt(val)}然后 5 条指令的Execute都变得只有一行:
func(self*ILOAD)Execute(frame*rtda.Frame){_iload(frame,uint(self.Index))}func(self*ILOAD_0)Execute(frame*rtda.Frame){_iload(frame,0)}func(self*ILOAD_1)Execute(frame*rtda.Frame){_iload(frame,1)}func(self*ILOAD_2)Execute(frame*rtda.Frame){_iload(frame,2)}func(self*ILOAD_3)Execute(frame*rtda.Frame){_iload(frame,3)}🎯这个模式会重复出现 10 次(5 种类型 × 2 个方向)。
_iload/_lload/_fload/_dload/_aload,以及_istore/_lstore/_fstore/_dstore/_astore。Go 里用下划线开头的函数名表示"包内私有",这是 Go 的惯例(虽然 Go 的可见性实际由首字母大小写决定,小写即私有)。
1.4 为什么是 0~3,而不是 0~7?
这是本篇最值得思考的设计问题。
局部变量索引使用频率(典型 Java 方法) slot 0 ████████████████████ ~35% ← 实例方法的 this / static 方法的第一个参数 slot 1 ██████████████ ~25% ← 第一个参数 slot 2 ██████ ~12% ← 第二个参数 slot 3 ███ ~6% ← 第三个参数 slot 4+ ████████ ~22% ← 走 iload <index>覆盖 78% 的场景,代价 4 个 opcode。如果扩到 0~7,覆盖率只提升到约 85%,却要多花 4 个 opcode——边际收益太低。
而且实际上,大部分方法的参数和局部变量总数不超过 4 个——《Clean Code》都建议方法参数不超过 3 个。
💡对比常量指令:int 常量给了 7 个 opcode(-1, 0~5),局部变量索引只给了 4 个(0~3)。因为索引的分布更集中(slot 0 就是 this 或首参,命中率极高)。
1.5 超过 255 个局部变量怎么办?
iload的索引是uint8,最大 255。如果方法的局部变量表超过 256 个槽位呢?
JVM 规范定义了wide指令来扩展——它会把索引扩展成 2 字节。这部分在第 040 篇讲。
普通形式: [0x15] [0x10] iload 16 (2 字节) wide 形式: [0xC4] [0x15] [0x01] [0x00] wide iload 256 (4 字节)实际上,超过 255 个局部变量的方法极其罕见(通常是代码生成器产出的),所以wide是一条"应急通道"。
二、存储指令:操作数栈 → 局部变量表
和加载指令刚好相反,存储指令把变量从操作数栈顶弹出,然后存入局部变量表。
2.1 33 条存储指令的全景
存储指令的 opcode 从0x36到0x56,同样是连续的 33 个,结构与加载指令完全对称:
opcode 助记符 说明 本章 ────────────────────────────────────────────────── 0x36 istore int,索引由操作数给出 ✅ 0x37 lstore long ✅ 0x38 fstore float ✅ 0x39 dstore double ✅ 0x3A astore reference ✅ ────────────────────────────────────────────────── 0x3B istore_0 int,索引隐含在操作码 ✅ 0x3C istore_1 ✅ 0x3D istore_2 ✅ 0x3E istore_3 ✅ 0x3F lstore_0 long ✅ 0x40 lstore_1 ✅ 0x41 lstore_2 ✅ 0x42 lstore_3 ✅ 0x43 fstore_0 float ✅ 0x44 fstore_1 ✅ 0x45 fstore_2 ✅ 0x46 fstore_3 ✅ 0x47 dstore_0 double ✅ 0x48 dstore_1 ✅ 0x49 dstore_2 ✅ 0x4A dstore_3 ✅ 0x4B astore_0 reference ✅ 0x4C astore_1 ✅ 0x4D astore_2 ✅ 0x4E astore_3 ✅ ────────────────────────────────────────────────── 0x4F iastore 数组元素存储(int) ❌ 第 8 章 0x50 lastore ❌ 0x51 fastore ❌ 0x52 dastore ❌ 0x53 aastore ❌ 0x54 bastore ❌ 0x55 castore ❌ 0x56 sastore ❌🔢记忆技巧:加载从
0x15开始,存储从0x36开始,两者相差0x21(33)。且istore的 opcode0x36=0x15 + 0x21。
2.2 代码实现:以 lstore 系列为例
在ch05/instructions/stores目录下创建lstore.go文件:
packagestoresimport"jvmgo/ch05/instructions/base"import"jvmgo/ch05/rtda"// Store long into local variabletypeLSTOREstruct{base.Index8Instruction}typeLSTORE_0struct{base.NoOperandsInstruction}typeLSTORE_1struct{base.NoOperandsInstruction}typeLSTORE_2struct{base.NoOperandsInstruction}typeLSTORE_3struct{base.NoOperandsInstruction}同样定义一个辅助函数供 5 条指令使用:
func_lstore(frame*rtda.Frame,indexuint){val:=frame.OperandStack().PopLong()frame.LocalVars().SetLong(index,val)}5 条指令的Execute:
func(self*LSTORE)Execute(frame*rtda.Frame){_lstore(frame,uint(self.Index))}func(self*LSTORE_0)Execute(frame*rtda.Frame){_lstore(frame,0)}func(self*LSTORE_1)Execute(frame*rtda.Frame){_lstore(frame,1)}func(self*LSTORE_2)Execute(frame*rtda.Frame){_lstore(frame,2)}func(self*LSTORE_3)Execute(frame*rtda.Frame){_lstore(frame,3)}2.3 加载 vs 存储的代码对称性
加载(loads) 存储(stores) ───────────────────────────── ───────────────────────────── func _iload(frame, index) { func _istore(frame, index) { val := frame val := frame .LocalVars() .OperandStack() .GetInt(index) .PopInt() frame frame .OperandStack() .LocalVars() .PushInt(val) .SetInt(index, val) } }只有两个区别:
- 数据来源 / 目的地对调(
LocalVars↔OperandStack) - 方法对调(
Get/Push↔Pop/Set)
三、long 和 double 的双槽位问题
第 4 章我们强调过:long和double在局部变量表和操作数栈中都占 2 个 Slot。这一点在加载/存储指令里体现得淋漓尽致。
3.1 局部变量表的双槽布局
publicstaticvoiddemo(){inta=1;// slot 0longb=2L;// slot 1, 2 ← 占两个!intc=3;// slot 3 ← 注意:不是 slot 2doubled=4.0;// slot 4, 5inte=5;// slot 6}局部变量表 LocalVars ┌─────────────────────────┐ │ slot 0 │ int a = 1 │ ├────────┼────────────────┤ │ slot 1 │ │ ← long b 的 │ slot 2 │ long b = 2L │ 两个槽位 ├────────┼────────────────┤ │ slot 3 │ int c = 3 │ ├────────┼────────────────┤ │ slot 4 │ │ ← double d 的 │ slot 5 │ double d = 4.0 │ 两个槽位 ├────────┼────────────────┤ │ slot 6 │ int e = 5 │ └────────┴────────────────┘对应的字节码:
0: iconst_1 1: istore_0 // a → slot 0 2: ldc2_w #2 // long 2l 5: lstore_1 // b → slot 1(占 1、2) 6: iconst_3 7: istore_3 // c → slot 3(不是 2!) 8: ldc2_w #4 // double 4.0d 11: dstore 4 // d → slot 4(占 4、5) 13: iconst_5 14: istore 6 // e → slot 6 16: return⚠️易错点:
long b存在 slot 1,那 slot 2 就不能再用了。下一个变量从 slot 3 开始。这是 JVM 规范明确规定的:“a value of type long or double occupies two consecutive local variables”。
3.2 为什么 long 和 double 要占两个槽?
因为 JVM 诞生时的目标是32 位机器。在 32 位架构上:
- 一个寄存器的宽度是 32 位
- 一个内存字(word)是 32 位
long(64 位)和double(64 位)必须拆成两半存
JVM 规范为了保持实现的简单性,直接规定:Slot 是 32 位的,long/double 占 2 个连续 Slot,低位在前。
long 0x0000000100000002L 存到 slot 1 和 slot 2: slot 1 = 0x00000002 ← 低 32 位 slot 2 = 0x00000001 ← 高 32 位这个约定叫"低位在前"(lower-order first),我们在第 027 篇实现SetLong/GetLong时已经处理过:
func(self LocalVars)SetLong(indexuint,valint64){self[index].num=int32(val)// 低 32 位放 indexself[index+1].num=int32(val>>32)// 高 32 位放 index+1}func(self LocalVars)GetLong(indexuint)int64{low:=uint32(self[index].num)high:=uint32(self[index+1].num)returnint64(high)<<32|int64(low)}3.3 对指令设计的影响
| 影响 | 说明 |
|---|---|
lstore_3会写 slot 3 和 4 | 所以max_locals必须 ≥ 5 |
lload_3会读 slot 3 和 4 | 不能只读一个 |
pop2/dup2的存在 | 因为pop只能弹 1 个 Slot,弹不了 long(第 035 篇) |
没有lload_4 | 如果用 slot 3,那就占 3 和 4,快捷指令的收益下降 |
四、实战:static 方法 vs 实例方法的局部变量表
这是理解aload_0的关键。
4.1 测试代码
publicclassSlotDemo{privateintvalue=10;// 实例方法publicintinstanceMethod(intx,inty){intsum=x+y;returnsum+value;}// static 方法publicstaticintstaticMethod(intx,inty){intsum=x+y;returnsum;}}4.2 反编译对比
javac SlotDemo.java javap-c-pSlotDemo.classinstanceMethod(实例方法):
public int instanceMethod(int, int); Code: 0: iload_1 // x 在 slot 1 1: iload_2 // y 在 slot 2 2: iadd 3: istore_3 // sum 在 slot 3 4: iload_3 5: aload_0 // ← this!slot 0 是 this 6: getfield #2 // Field value:I 9: iadd 10: ireturnstaticMethod(静态方法):
public static int staticMethod(int, int); Code: 0: iload_0 // x 在 slot 0 ← 没有 this! 1: iload_1 // y 在 slot 1 2: iadd 3: istore_2 // sum 在 slot 2 4: iload_2 5: ireturn4.3 局部变量表布局对比图
instanceMethod(x, y) staticMethod(x, y) ┌────────────────────┐ ┌────────────────────┐ │ slot 0: this │ │ slot 0: x │ │ slot 1: x │ │ slot 1: y │ │ slot 2: y │ │ slot 2: sum │ │ slot 3: sum │ └────────────────────┘ └────────────────────┘ ↑ ↑ this 占用 slot 0 没有 this, 参数从 slot 1 开始 参数从 slot 0 开始结论:
| 方法类型 | slot 0 | 参数起始 | 访问实例字段 |
|---|---|---|---|
| 实例方法 | this(reference) | slot 1 | 必须先aload_0,再getfield |
| static 方法 | 第一个参数 | slot 0 | 只能getstatic |
💡这就是 Java 里
this的本质:它不是关键字,而是编译器自动加到实例方法里的第 0 号局部变量。访问this.value编译后就是aload_0+getfield #2。这个设计还解释了一个经典问题:为什么 static 方法里不能用
this?因为 static 方法的局部变量表里根本没有 slot 0 给 this,编译器无从生成aload_0。
4.4 另一个经典问题:++ 的字节码
publicvoidinc(intx){x++;// 编译成什么?}public void inc(int); Code: 0: iinc 1, 1 ← 直接改局部变量表! 3: return注意:i++作为独立语句时,编译器用的是iinc指令,而不是iload+iconst_1+iadd+istore。
iinc直接操作局部变量表,完全不经过操作数栈:
iload/iconst/iadd/istore 路径: iinc 路径: 局部变量表 ──iload──► 操作数栈 局部变量表 ▲ │ │ │ ▼ ▼ └──istore────── iadd iinc(原地 +1) ▲ │ iconst_1 4 条指令,2 次栈操作 1 条指令,0 次栈操作iinc属于数学指令,我们在第 036 篇讲它。
五、loads 和 stores 包的文件清单
ch05/instructions/loads/ ch05/instructions/stores/ ├── iload.go (ILOAD, ILOAD_0~3) ├── istore.go (ISTORE, ISTORE_0~3) ├── lload.go (LLOAD, LLOAD_0~3) ├── lstore.go (LSTORE, LSTORE_0~3) ├── fload.go (FLOAD, FLOAD_0~3) ├── fstore.go (FSTORE, FSTORE_0~3) ├── dload.go (DLOAD, DLOAD_0~3) ├── dstore.go (DSTORE, DSTORE_0~3) └── aload.go (ALOAD, ALOAD_0~3) └── astore.go (ASTORE, ASTORE_0~3) ───────────────────────────────── ───────────────────────────────── 5 个文件,25 条指令 5 个文件,25 条指令总共 10 个文件、50 条指令。每个文件的结构完全一致,以aload.go为例:
packageloadsimport"jvmgo/ch05/instructions/base"import"jvmgo/ch05/rtda"// Load reference from local variabletypeALOADstruct{base.Index8Instruction}typeALOAD_0struct{base.NoOperandsInstruction}typeALOAD_1struct{base.NoOperandsInstruction}typeALOAD_2struct{base.NoOperandsInstruction}typeALOAD_3struct{base.NoOperandsInstruction}func_aload(frame*rtda.Frame,indexuint){ref:=frame.LocalVars().GetRef(index)frame.OperandStack().PushRef(ref)}func(self*ALOAD)Execute(frame*rtda.Frame){_aload(frame,uint(self.Index))}func(self*ALOAD_0)Execute(frame*rtda.Frame){_aload(frame,0)}func(self*ALOAD_1)Execute(frame*rtda.Frame){_aload(frame,1)}func(self*ALOAD_2)Execute(frame*rtda.Frame){_aload(frame,2)}func(self*ALOAD_3)Execute(frame*rtda.Frame){_aload(frame,3)}对应的astore.go:
packagestoresimport"jvmgo/ch05/instructions/base"import"jvmgo/ch05/rtda"// Store reference into local variabletypeASTOREstruct{base.Index8Instruction}typeASTORE_0struct{base.NoOperandsInstruction}typeASTORE_1struct{base.NoOperandsInstruction}typeASTORE_2struct{base.NoOperandsInstruction}typeASTORE_3struct{base.NoOperandsInstruction}func_astore(frame*rtda.Frame,indexuint){ref:=frame.OperandStack().PopRef()frame.LocalVars().SetRef(index,ref)}func(self*ASTORE)Execute(frame*rtda.Frame){_astore(frame,uint(self.Index))}// ASTORE_0 ~ ASTORE_3 同理,_astore(frame, 0~3)🎯10 个文件的代码结构几乎一模一样——这正是"用模式消除重复"的价值。写第一个文件时会觉得啰嗦,后面 9 个就是纯粹的复制 + 改类型名了。
六、单元测试
packagestoresimport("testing""jvmgo/ch05/rtda")funcTestStoreAndLoad(t*testing.T){frame:=rtda.NewFrame(nil,10,10)localVars:=frame.LocalVars()stack:=frame.OperandStack()// 验证 long 占两个 slotstack.PushLong(0x0000000100000002)(&LSTORE_1{}).Execute(frame)// 存到 slot 1 和 2iflocalVars.GetInt(3)!=0{t.Errorf("slot 3 should be untouched")}// 读回来(&LLOAD_1{}).Execute(frame)ifgot:=stack.PopLong();got!=0x0000000100000002{t.Errorf("lload_1: got 0x%x",got)}// 验证操作数栈 size 变化:long 占 2 个 slotifstack.Size()!=0{t.Errorf("stack should be empty, got %d",stack.Size())}stack.PushLong(1)ifstack.Size()!=2{t.Errorf("long should occupy 2 slots, got %d",stack.Size())}}🧪这个测试覆盖了本篇最容易出错的两点:
lstore_1写 slot 1 和 2,不能影响 slot 3PushLong让操作数栈的size加 2,不是加 1
本篇小结
加载指令和存储指令是 JVM 中数据的搬运工,本篇实现了 50 条:
- 加载指令(loads)25 条——
iload/lload/fload/dload/aload及其_0~_3快捷形式,把局部变量表的数据推入操作数栈。 - 存储指令(stores)25 条——
istore/lstore/fstore/dstore/astore及其_0~_3快捷形式,把栈顶数据弹回局部变量表。 - 数组相关的 8 + 8 条(
xaload/xastore)依赖数组对象和堆内存,留到第 8 章。
四条核心规律:
- opcode 布局完美对称:加载
0x15~0x35,存储0x36~0x56,相差0x21(33)。 - 快捷形式只有
_0~_3:覆盖约 78% 的实际使用,继续扩展的边际收益太低。 - long/double 占两个 Slot:
lstore_1写 slot 1 和 2,下一个变量从 slot 3 开始;PushLong让栈 size +2。 - 实例方法的 slot 0 是
this:这就是this的本质——编译器自动加的第 0 号局部变量。static 方法没有 this,参数从 slot 0 开始。
下一篇进入栈指令——pop/dup/swap这一组直接操作操作数栈的指令。它们虽然只有 9 条,却是整个指令集里最烧脑的一类,尤其是dup2_x2这类带位插入的复制指令。
上一篇【第33篇】常量指令——把数字推到栈上的艺术
下一篇【第35篇】栈指令——dup/swap 的奇妙世界