☰
【DIY系列:Java虚拟机】第34篇:加载与存储指令——局部变量表和操作数栈的“桥梁“
2026/9/29 3:40:15 网站建设 项目流程

上一篇【第33篇】常量指令——把数字推到栈上的艺术
下一篇【第35篇】栈指令——dup/swap 的奇妙世界


摘要

上一篇我们造了常量指令——把值无中生有地推到操作数栈上。但 JVM 里的数据还有一个更重要的来源:局部变量。

这一篇要造的是加载指令(loads)和存储指令(stores),它们是局部变量表和操作数栈之间的唯一桥梁:

局部变量表 操作数栈 LocalVars xload OperandStack ┌────────┐ ────────► ┌────────┐ │ slot 0 │ │ │ │ slot 1 │ ◄──────── │ 栈顶 │ │ slot 2 │ xstore │ │ └────────┘ └────────┘

JVM 规范给这两类各分配了33 条指令,合计 66 条——占全部 205 条指令的32%。本章实现其中的 50 条(数组相关的xaload/xastore留到第 8 章)。

读完这一篇,你会理解:为什么局部变量索引只有 0~3 有快捷指令?为什么iload_3存在而iload_4不存在?超过 255 个局部变量怎么办?


一、加载指令:局部变量表 → 操作数栈

1.1 33 条加载指令的全景

加载指令从局部变量表获取变量,然后推入操作数栈顶。按照所操作变量的类型可以分为6 类,opcode 从0x15到0x35,恰好连续 33 个:

opcode 助记符 说明 本章 ────────────────────────────────────────────────── 0x15 iload int,索引由操作数给出 ✅ 0x16 lload long ✅ 0x17 fload float ✅ 0x18 dload double ✅ 0x19 aload reference ✅ ────────────────────────────────────────────────── 0x1A iload_0 int,索引隐含在操作码 ✅ 0x1B iload_1 ✅ 0x1C iload_2 ✅ 0x1D iload_3 ✅ 0x1E lload_0 long ✅ 0x1F lload_1 ✅ 0x20 lload_2 ✅ 0x21 lload_3 ✅ 0x22 fload_0 float ✅ 0x23 fload_1 ✅ 0x24 fload_2 ✅ 0x25 fload_3 ✅ 0x26 dload_0 double ✅ 0x27 dload_1 ✅ 0x28 dload_2 ✅ 0x29 dload_3 ✅ 0x2A aload_0 reference ✅ 0x2B aload_1 ✅ 0x2C aload_2 ✅ 0x2D aload_3 ✅ ────────────────────────────────────────────────── 0x2E iaload 数组元素加载(int) ❌ 第 8 章 0x2F laload ❌ 0x30 faload ❌ 0x31 daload ❌ 0x32 aaload ❌ 0x33 baload ❌ 0x34 caload ❌ 0x35 saload ❌

规律极其整齐:

类型带索引版快捷版数组版
intiloadiload_0~iload_3iaload
longlloadlload_0~lload_3laload
floatfloadfload_0~fload_3faload
doubledloaddload_0~dload_3daload
referencealoadaload_0~aload_3aaload
byte/char/short——baload/caload/saload

5 × (1 + 4) + 8 =33✅

1.2 代码实现:以 iload 系列为例

在ch05/instructions/loads目录下创建iload.go文件:

packageloadsimport"jvmgo/ch05/instructions/base"import"jvmgo/ch05/rtda"// Load int from local variabletypeILOADstruct{base.Index8Instruction}typeILOAD_0struct{base.NoOperandsInstruction}typeILOAD_1struct{base.NoOperandsInstruction}typeILOAD_2struct{base.NoOperandsInstruction}typeILOAD_3struct{base.NoOperandsInstruction}

注意这里两种抽象指令的混合使用:

  • ILOAD嵌入base.Index8Instruction——索引来自操作数(1 字节),需要FetchOperands去读
  • ILOAD_0~ILOAD_3嵌入base.NoOperandsInstruction——索引写死在类型里,没有操作数

1.3 用辅助函数消除重复

为了避免重复代码,定义一个函数供 5 条指令共用:

func_iload(frame*rtda.Frame,indexuint){val:=frame.LocalVars().GetInt(index)frame.OperandStack().PushInt(val)}

然后 5 条指令的Execute都变得只有一行:

func(self*ILOAD)Execute(frame*rtda.Frame){_iload(frame,uint(self.Index))}func(self*ILOAD_0)Execute(frame*rtda.Frame){_iload(frame,0)}func(self*ILOAD_1)Execute(frame*rtda.Frame){_iload(frame,1)}func(self*ILOAD_2)Execute(frame*rtda.Frame){_iload(frame,2)}func(self*ILOAD_3)Execute(frame*rtda.Frame){_iload(frame,3)}

🎯这个模式会重复出现 10 次(5 种类型 × 2 个方向)。_iload/_lload/_fload/_dload/_aload,以及_istore/_lstore/_fstore/_dstore/_astore。

Go 里用下划线开头的函数名表示"包内私有",这是 Go 的惯例(虽然 Go 的可见性实际由首字母大小写决定,小写即私有)。

1.4 为什么是 0~3,而不是 0~7?

这是本篇最值得思考的设计问题。

局部变量索引使用频率(典型 Java 方法) slot 0 ████████████████████ ~35% ← 实例方法的 this / static 方法的第一个参数 slot 1 ██████████████ ~25% ← 第一个参数 slot 2 ██████ ~12% ← 第二个参数 slot 3 ███ ~6% ← 第三个参数 slot 4+ ████████ ~22% ← 走 iload <index>

覆盖 78% 的场景,代价 4 个 opcode。如果扩到 0~7,覆盖率只提升到约 85%,却要多花 4 个 opcode——边际收益太低。

而且实际上,大部分方法的参数和局部变量总数不超过 4 个——《Clean Code》都建议方法参数不超过 3 个。

💡对比常量指令:int 常量给了 7 个 opcode(-1, 0~5),局部变量索引只给了 4 个(0~3)。因为索引的分布更集中(slot 0 就是 this 或首参,命中率极高)。

1.5 超过 255 个局部变量怎么办?

iload的索引是uint8,最大 255。如果方法的局部变量表超过 256 个槽位呢?

JVM 规范定义了wide指令来扩展——它会把索引扩展成 2 字节。这部分在第 040 篇讲。

普通形式: [0x15] [0x10] iload 16 (2 字节) wide 形式: [0xC4] [0x15] [0x01] [0x00] wide iload 256 (4 字节)

实际上,超过 255 个局部变量的方法极其罕见(通常是代码生成器产出的),所以wide是一条"应急通道"。


二、存储指令:操作数栈 → 局部变量表

和加载指令刚好相反,存储指令把变量从操作数栈顶弹出,然后存入局部变量表。

2.1 33 条存储指令的全景

存储指令的 opcode 从0x36到0x56,同样是连续的 33 个,结构与加载指令完全对称:

opcode 助记符 说明 本章 ────────────────────────────────────────────────── 0x36 istore int,索引由操作数给出 ✅ 0x37 lstore long ✅ 0x38 fstore float ✅ 0x39 dstore double ✅ 0x3A astore reference ✅ ────────────────────────────────────────────────── 0x3B istore_0 int,索引隐含在操作码 ✅ 0x3C istore_1 ✅ 0x3D istore_2 ✅ 0x3E istore_3 ✅ 0x3F lstore_0 long ✅ 0x40 lstore_1 ✅ 0x41 lstore_2 ✅ 0x42 lstore_3 ✅ 0x43 fstore_0 float ✅ 0x44 fstore_1 ✅ 0x45 fstore_2 ✅ 0x46 fstore_3 ✅ 0x47 dstore_0 double ✅ 0x48 dstore_1 ✅ 0x49 dstore_2 ✅ 0x4A dstore_3 ✅ 0x4B astore_0 reference ✅ 0x4C astore_1 ✅ 0x4D astore_2 ✅ 0x4E astore_3 ✅ ────────────────────────────────────────────────── 0x4F iastore 数组元素存储(int) ❌ 第 8 章 0x50 lastore ❌ 0x51 fastore ❌ 0x52 dastore ❌ 0x53 aastore ❌ 0x54 bastore ❌ 0x55 castore ❌ 0x56 sastore ❌

🔢记忆技巧:加载从0x15开始,存储从0x36开始,两者相差0x21(33)。且istore的 opcode0x36=0x15 + 0x21。

2.2 代码实现:以 lstore 系列为例

在ch05/instructions/stores目录下创建lstore.go文件:

packagestoresimport"jvmgo/ch05/instructions/base"import"jvmgo/ch05/rtda"// Store long into local variabletypeLSTOREstruct{base.Index8Instruction}typeLSTORE_0struct{base.NoOperandsInstruction}typeLSTORE_1struct{base.NoOperandsInstruction}typeLSTORE_2struct{base.NoOperandsInstruction}typeLSTORE_3struct{base.NoOperandsInstruction}

同样定义一个辅助函数供 5 条指令使用:

func_lstore(frame*rtda.Frame,indexuint){val:=frame.OperandStack().PopLong()frame.LocalVars().SetLong(index,val)}

5 条指令的Execute:

func(self*LSTORE)Execute(frame*rtda.Frame){_lstore(frame,uint(self.Index))}func(self*LSTORE_0)Execute(frame*rtda.Frame){_lstore(frame,0)}func(self*LSTORE_1)Execute(frame*rtda.Frame){_lstore(frame,1)}func(self*LSTORE_2)Execute(frame*rtda.Frame){_lstore(frame,2)}func(self*LSTORE_3)Execute(frame*rtda.Frame){_lstore(frame,3)}

2.3 加载 vs 存储的代码对称性

加载(loads) 存储(stores) ───────────────────────────── ───────────────────────────── func _iload(frame, index) { func _istore(frame, index) { val := frame val := frame .LocalVars() .OperandStack() .GetInt(index) .PopInt() frame frame .OperandStack() .LocalVars() .PushInt(val) .SetInt(index, val) } }

只有两个区别:

  1. 数据来源 / 目的地对调(LocalVars↔OperandStack)
  2. 方法对调(Get/Push↔Pop/Set)

三、long 和 double 的双槽位问题

第 4 章我们强调过:long和double在局部变量表和操作数栈中都占 2 个 Slot。这一点在加载/存储指令里体现得淋漓尽致。

3.1 局部变量表的双槽布局

publicstaticvoiddemo(){inta=1;// slot 0longb=2L;// slot 1, 2 ← 占两个!intc=3;// slot 3 ← 注意:不是 slot 2doubled=4.0;// slot 4, 5inte=5;// slot 6}
局部变量表 LocalVars ┌─────────────────────────┐ │ slot 0 │ int a = 1 │ ├────────┼────────────────┤ │ slot 1 │ │ ← long b 的 │ slot 2 │ long b = 2L │ 两个槽位 ├────────┼────────────────┤ │ slot 3 │ int c = 3 │ ├────────┼────────────────┤ │ slot 4 │ │ ← double d 的 │ slot 5 │ double d = 4.0 │ 两个槽位 ├────────┼────────────────┤ │ slot 6 │ int e = 5 │ └────────┴────────────────┘

对应的字节码:

0: iconst_1 1: istore_0 // a → slot 0 2: ldc2_w #2 // long 2l 5: lstore_1 // b → slot 1(占 1、2) 6: iconst_3 7: istore_3 // c → slot 3(不是 2!) 8: ldc2_w #4 // double 4.0d 11: dstore 4 // d → slot 4(占 4、5) 13: iconst_5 14: istore 6 // e → slot 6 16: return

⚠️易错点:long b存在 slot 1,那 slot 2 就不能再用了。下一个变量从 slot 3 开始。这是 JVM 规范明确规定的:“a value of type long or double occupies two consecutive local variables”。

3.2 为什么 long 和 double 要占两个槽?

因为 JVM 诞生时的目标是32 位机器。在 32 位架构上:

  • 一个寄存器的宽度是 32 位
  • 一个内存字(word)是 32 位
  • long(64 位)和double(64 位)必须拆成两半存

JVM 规范为了保持实现的简单性,直接规定:Slot 是 32 位的,long/double 占 2 个连续 Slot,低位在前。

long 0x0000000100000002L 存到 slot 1 和 slot 2: slot 1 = 0x00000002 ← 低 32 位 slot 2 = 0x00000001 ← 高 32 位

这个约定叫"低位在前"(lower-order first),我们在第 027 篇实现SetLong/GetLong时已经处理过:

func(self LocalVars)SetLong(indexuint,valint64){self[index].num=int32(val)// 低 32 位放 indexself[index+1].num=int32(val>>32)// 高 32 位放 index+1}func(self LocalVars)GetLong(indexuint)int64{low:=uint32(self[index].num)high:=uint32(self[index+1].num)returnint64(high)<<32|int64(low)}

3.3 对指令设计的影响

影响说明
lstore_3会写 slot 3 和 4所以max_locals必须 ≥ 5
lload_3会读 slot 3 和 4不能只读一个
pop2/dup2的存在因为pop只能弹 1 个 Slot,弹不了 long(第 035 篇)
没有lload_4如果用 slot 3,那就占 3 和 4,快捷指令的收益下降

四、实战:static 方法 vs 实例方法的局部变量表

这是理解aload_0的关键。

4.1 测试代码

publicclassSlotDemo{privateintvalue=10;// 实例方法publicintinstanceMethod(intx,inty){intsum=x+y;returnsum+value;}// static 方法publicstaticintstaticMethod(intx,inty){intsum=x+y;returnsum;}}

4.2 反编译对比

javac SlotDemo.java javap-c-pSlotDemo.class

instanceMethod(实例方法):

public int instanceMethod(int, int); Code: 0: iload_1 // x 在 slot 1 1: iload_2 // y 在 slot 2 2: iadd 3: istore_3 // sum 在 slot 3 4: iload_3 5: aload_0 // ← this!slot 0 是 this 6: getfield #2 // Field value:I 9: iadd 10: ireturn

staticMethod(静态方法):

public static int staticMethod(int, int); Code: 0: iload_0 // x 在 slot 0 ← 没有 this! 1: iload_1 // y 在 slot 1 2: iadd 3: istore_2 // sum 在 slot 2 4: iload_2 5: ireturn

4.3 局部变量表布局对比图

instanceMethod(x, y) staticMethod(x, y) ┌────────────────────┐ ┌────────────────────┐ │ slot 0: this │ │ slot 0: x │ │ slot 1: x │ │ slot 1: y │ │ slot 2: y │ │ slot 2: sum │ │ slot 3: sum │ └────────────────────┘ └────────────────────┘ ↑ ↑ this 占用 slot 0 没有 this, 参数从 slot 1 开始 参数从 slot 0 开始

结论:

方法类型slot 0参数起始访问实例字段
实例方法this(reference)slot 1必须先aload_0,再getfield
static 方法第一个参数slot 0只能getstatic

💡这就是 Java 里this的本质:它不是关键字,而是编译器自动加到实例方法里的第 0 号局部变量。访问this.value编译后就是aload_0+getfield #2。

这个设计还解释了一个经典问题:为什么 static 方法里不能用this?因为 static 方法的局部变量表里根本没有 slot 0 给 this,编译器无从生成aload_0。

4.4 另一个经典问题:++ 的字节码

publicvoidinc(intx){x++;// 编译成什么?}
public void inc(int); Code: 0: iinc 1, 1 ← 直接改局部变量表! 3: return

注意:i++作为独立语句时,编译器用的是iinc指令,而不是iload+iconst_1+iadd+istore。

iinc直接操作局部变量表,完全不经过操作数栈:

iload/iconst/iadd/istore 路径: iinc 路径: 局部变量表 ──iload──► 操作数栈 局部变量表 ▲ │ │ │ ▼ ▼ └──istore────── iadd iinc(原地 +1) ▲ │ iconst_1 4 条指令,2 次栈操作 1 条指令,0 次栈操作

iinc属于数学指令,我们在第 036 篇讲它。


五、loads 和 stores 包的文件清单

ch05/instructions/loads/ ch05/instructions/stores/ ├── iload.go (ILOAD, ILOAD_0~3) ├── istore.go (ISTORE, ISTORE_0~3) ├── lload.go (LLOAD, LLOAD_0~3) ├── lstore.go (LSTORE, LSTORE_0~3) ├── fload.go (FLOAD, FLOAD_0~3) ├── fstore.go (FSTORE, FSTORE_0~3) ├── dload.go (DLOAD, DLOAD_0~3) ├── dstore.go (DSTORE, DSTORE_0~3) └── aload.go (ALOAD, ALOAD_0~3) └── astore.go (ASTORE, ASTORE_0~3) ───────────────────────────────── ───────────────────────────────── 5 个文件,25 条指令 5 个文件,25 条指令

总共 10 个文件、50 条指令。每个文件的结构完全一致,以aload.go为例:

packageloadsimport"jvmgo/ch05/instructions/base"import"jvmgo/ch05/rtda"// Load reference from local variabletypeALOADstruct{base.Index8Instruction}typeALOAD_0struct{base.NoOperandsInstruction}typeALOAD_1struct{base.NoOperandsInstruction}typeALOAD_2struct{base.NoOperandsInstruction}typeALOAD_3struct{base.NoOperandsInstruction}func_aload(frame*rtda.Frame,indexuint){ref:=frame.LocalVars().GetRef(index)frame.OperandStack().PushRef(ref)}func(self*ALOAD)Execute(frame*rtda.Frame){_aload(frame,uint(self.Index))}func(self*ALOAD_0)Execute(frame*rtda.Frame){_aload(frame,0)}func(self*ALOAD_1)Execute(frame*rtda.Frame){_aload(frame,1)}func(self*ALOAD_2)Execute(frame*rtda.Frame){_aload(frame,2)}func(self*ALOAD_3)Execute(frame*rtda.Frame){_aload(frame,3)}

对应的astore.go:

packagestoresimport"jvmgo/ch05/instructions/base"import"jvmgo/ch05/rtda"// Store reference into local variabletypeASTOREstruct{base.Index8Instruction}typeASTORE_0struct{base.NoOperandsInstruction}typeASTORE_1struct{base.NoOperandsInstruction}typeASTORE_2struct{base.NoOperandsInstruction}typeASTORE_3struct{base.NoOperandsInstruction}func_astore(frame*rtda.Frame,indexuint){ref:=frame.OperandStack().PopRef()frame.LocalVars().SetRef(index,ref)}func(self*ASTORE)Execute(frame*rtda.Frame){_astore(frame,uint(self.Index))}// ASTORE_0 ~ ASTORE_3 同理,_astore(frame, 0~3)

🎯10 个文件的代码结构几乎一模一样——这正是"用模式消除重复"的价值。写第一个文件时会觉得啰嗦,后面 9 个就是纯粹的复制 + 改类型名了。


六、单元测试

packagestoresimport("testing""jvmgo/ch05/rtda")funcTestStoreAndLoad(t*testing.T){frame:=rtda.NewFrame(nil,10,10)localVars:=frame.LocalVars()stack:=frame.OperandStack()// 验证 long 占两个 slotstack.PushLong(0x0000000100000002)(&LSTORE_1{}).Execute(frame)// 存到 slot 1 和 2iflocalVars.GetInt(3)!=0{t.Errorf("slot 3 should be untouched")}// 读回来(&LLOAD_1{}).Execute(frame)ifgot:=stack.PopLong();got!=0x0000000100000002{t.Errorf("lload_1: got 0x%x",got)}// 验证操作数栈 size 变化:long 占 2 个 slotifstack.Size()!=0{t.Errorf("stack should be empty, got %d",stack.Size())}stack.PushLong(1)ifstack.Size()!=2{t.Errorf("long should occupy 2 slots, got %d",stack.Size())}}

🧪这个测试覆盖了本篇最容易出错的两点:

  1. lstore_1写 slot 1 和 2,不能影响 slot 3
  2. PushLong让操作数栈的size加 2,不是加 1

本篇小结

加载指令和存储指令是 JVM 中数据的搬运工,本篇实现了 50 条:

  1. 加载指令(loads)25 条——iload/lload/fload/dload/aload及其_0~_3快捷形式,把局部变量表的数据推入操作数栈。
  2. 存储指令(stores)25 条——istore/lstore/fstore/dstore/astore及其_0~_3快捷形式,把栈顶数据弹回局部变量表。
  3. 数组相关的 8 + 8 条(xaload/xastore)依赖数组对象和堆内存,留到第 8 章。

四条核心规律:

  • opcode 布局完美对称:加载0x15~0x35,存储0x36~0x56,相差0x21(33)。
  • 快捷形式只有_0~_3:覆盖约 78% 的实际使用,继续扩展的边际收益太低。
  • long/double 占两个 Slot:lstore_1写 slot 1 和 2,下一个变量从 slot 3 开始;PushLong让栈 size +2。
  • 实例方法的 slot 0 是this:这就是this的本质——编译器自动加的第 0 号局部变量。static 方法没有 this,参数从 slot 0 开始。

下一篇进入栈指令——pop/dup/swap这一组直接操作操作数栈的指令。它们虽然只有 9 条,却是整个指令集里最烧脑的一类,尤其是dup2_x2这类带位插入的复制指令。


上一篇【第33篇】常量指令——把数字推到栈上的艺术
下一篇【第35篇】栈指令——dup/swap 的奇妙世界


需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询