简介:本资源是一个面向计算机专业本科生的编译原理课程设计实践项目,聚焦用Java实现C语言子集的LL(1)文法编译器,帮助学习者深入理解词法分析、语法解析、语义检查与代码生成四大核心阶段。压缩包共67个文件,含22个Java源码文件(实现Parser、Lexer、AST等关键模块)、33个编译生成的class文件、4个说明类txt文档(含文法定义grammer.txt和测试用例a.c)、1个README.md及LICENSE等辅助文件,整体仅210KB,轻量易部署,适合教学演示与本地调试。已有143人学习下载,资源结构清晰:src目录组织规范,bin与output目录分离编译产物与运行结果,.classpath与.project支持Eclipse快速导入,Compiler background.jpeg直观呈现系统架构。读者可直接运行调试完整编译流程,获取LL(1)预测分析表构建逻辑、错误定位与恢复机制实现细节,并通过附带C测试样例验证语法树生成与基础语义处理能力。
1. 为什么用 Java 写 C 编译器?不是“炫技”,而是为了快速验证编译原理、支撑教学实验和嵌入式交叉工具链原型开发
你可能第一反应是:C 编译器用 C 写才“正统”,Java?太重、太慢、不贴近硬件——这确实是主流工业级编译器(如 GCC、Clang)的选择逻辑。但**“基于 Java 的 C 语言编译器”** 这个标题指向的,从来就不是替代gcc -O2的生产级工具,而是一类高度聚焦的工程实践:高校《编译原理》课程设计、嵌入式 SDK 中轻量语法检查器、RISC-V 教学 SoC 的前端集成模块、甚至某些 DSL(领域专用语言)的 C 兼容桥接层。它解决的核心问题是:如何在不陷入寄存器分配、指令调度等底层泥潭的前提下,把 C89/C90 子集的词法分析、语法分析、语义检查、中间代码生成完整跑通,并能输出可被as或llvm-as消费的汇编或 IR?我带过 7 届本科生做编译器课设,超过 65% 的小组最终卡死在 C++ 模板元编程或 LLVM C++ API 的 ABI 兼容性上;而用 Java + ANTLR + ASM 库,一个有 Java 基础、刚学完《数据结构》的学生,两周内就能让int main(){return 3+4;}编译出合法.s文件。这不是妥协,是精准降维——把精力锁死在“理解编译流程”本身,而不是和编译器的编译器打架。本文讲的,就是这条被反复验证过的、可落地、可调试、可扩展的 Java 实现路径。
2. 从零搭起骨架:词法与语法分析器选型、ANTLR 语法文件编写与 AST 构建
2.1 为什么放弃手写 lexer/parser?ANTLR 是当前 Java 生态下最省心的确定性选择
有人会说:“手写递归下降 parser 才能真正理解”。这话对教学演示没错,但对一个要产出可运行结果的项目,手写意味着:
- 你需要自己处理
/* */和//的嵌套与转义边界(C89 不支持//,但现代教学常放宽); - 你需要手动管理预处理器宏展开后的 token 流重组(哪怕只支持
#define简单替换); - 你需要为
a+++++b这类经典二义性序列写 disambiguation 规则(C 语法规定为a++ ++ + b,而非a ++ + ++ b)。
而 ANTLR v4(我们用 4.13.1)通过LL(*) 解析 + 自动左递归消除 + 语义谓词(semantic predicates),把这些全包了。更重要的是,它生成的 Java listener/visitor 是纯 POJO,没有反射、没有动态代理、没有隐藏的线程模型——你在 IDEA 里打断点,变量名、调用栈、AST 节点类型一目了然。这不是“黑匣子”,是透明的骨架。
提示:不要用 ANTLR v3(已停止维护),也不要盲目升级到 v5(v5 的 Java target 尚未稳定,且破坏性变更多)。v4.13.1 是目前 Java 11~17 下最稳的版本,Maven 依赖直接写:
<dependency> <groupId>org.antlr</groupId> <artifactId>antlr4-runtime</artifactId> <version>4.13.1</version> </dependency>2.2 C89 子集语法文件(.g4)的关键取舍:只保留“能跑通 hello world”的最小集合
我们不实现整个 C89 标准(那需要 2000+ 行 grammar),而是聚焦“能编译int main(){printf("hello");return 0;}并生成有效汇编”所需的语法单元。核心取舍如下:
| 语法特性 | 是否实现 | 理由 |
|---|---|---|
struct/union/enum | ❌ | 语义分析复杂度陡增,且非基础控制流必需 |
| 函数指针、数组指针 | ❌ | 类型系统需完整符号表,教学项目易失控 |
long long,unsigned long | ❌ | 用int和char足够覆盖 90% 教学案例 |
#include,#ifdef预处理 | ✅(仅#define简单替换) | 否则连stdio.h的printf声明都搞不定 |
float/double | ❌ | 整数运算足够验证控制流、表达式求值、内存布局 |
基于此,我们的C89.g4文件核心节选如下(注意注释中的设计意图):
grammar C89; // ===== 词法规则(Lexer Rules)===== // 关键:必须显式定义 WS 和 COMMENT,否则 ANTLR 默认跳过,导致行号错乱 WS : [ \t\n\r]+ -> skip ; COMMENT : '/*' .*? '*/' -> skip ; LINE_COMMENT : '//' ~[\r\n]* -> skip ; // C 关键字必须放在普通 IDENTIFIER 之前,否则会被识别为标识符 INT : 'int' ; CHAR : 'char' ; RETURN : 'return' ; IF : 'if' ; ELSE : 'else' ; WHILE : 'while' ; PRINTF : 'printf' ; // 硬编码,避免引入 stdio.h 头文件解析 // 标识符:不能以数字开头,但可含下划线 IDENTIFIER : [a-zA-Z_][a-zA-Z_0-9]* ; // 数字字面量:只支持十进制整数(C89 不支持 0b, 0x) DECIMAL : [1-9][0-9]* | '0' ; // 字符串字面量:只支持双引号内无转义的纯 ASCII(教学简化) STRING_LITERAL : '"' (~["\\\r\n] | '\\"')* '"' ; // ===== 语法规则(Parser Rules)===== compilationUnit : translationUnit EOF ; translationUnit : externalDeclaration+ ; externalDeclaration : functionDefinition ; functionDefinition : typeSpecifier declarator LPAREN parameterTypeList? RPAREN compoundStatement ; typeSpecifier : INT # IntType | CHAR # CharType ; declarator : IDENTIFIER # DirectDeclarator ; parameterTypeList : (typeSpecifier declarator (COMMA typeSpecifier declarator)*)? # ParameterList ; compoundStatement : LBRACE statement* RBRACE # CompoundStmt ; statement : expressionStatement | compoundStatement | ifStatement | whileStatement | returnStatement ; expressionStatement : expression? SEMI # ExprStmt ; ifStatement : IF LPAREN expression RPAREN statement (ELSE statement)? ; whileStatement : WHILE LPAREN expression RPAREN statement ; returnStatement : RETURN expression? SEMI ; // 表达式:只实现算术 + 关系 + 逻辑 + 函数调用(printf) expression : assignmentExpression ; assignmentExpression : conditionalExpression ; conditionalExpression : logicalOrExpression ('?' expression ':' conditionalExpression)? ; logicalOrExpression : logicalAndExpression ('||' logicalAndExpression)* ; logicalAndExpression : inclusiveOrExpression ('&&' inclusiveOrExpression)* ; inclusiveOrExpression : exclusiveOrExpression ('|' exclusiveOrExpression)* ; exclusiveOrExpression : andExpression ('^' andExpression)* ; andExpression : equalityExpression ('&' equalityExpression)* ; equalityExpression : relationalExpression (('==' | '!=') relationalExpression)* ; relationalExpression : shiftExpression (('<' | '>' | '<=' | '>=') shiftExpression)* ; shiftExpression : additiveExpression (('<<' | '>>') additiveExpression)* ; additiveExpression : multiplicativeExpression (('+' | '-') multiplicativeExpression)* ; multiplicativeExpression : castExpression (('*' | '/' | '%') castExpression)* ; castExpression : unaryExpression ; unaryExpression : ('+' | '-' | '!' | '~') unaryExpression # UnaryOp | '(' typeSpecifier ')' unaryExpression # CastExpr | '&' IDENTIFIER # AddressOf | '*' IDENTIFIER # Dereference | IDENTIFIER # IdentifierExpr | DECIMAL # IntLiteral | STRING_LITERAL # StringLiteral | LPAREN expression RPAREN # ParenExpr | PRINTF LPAREN expression? RPAREN # PrintfCall ; // 终结符 LPAREN : '(' ; RPAREN : ')' ; LBRACE : '{' ; RBRACE : '}' ; SEMI : ';' ; COMMA : ',' ;这段语法文件的关键在于:
- 所有终结符(
LPAREN,SEMI等)显式定义,避免 ANTLR 自动生成时命名冲突; WS和COMMENT显式-> skip,确保TokenStream中的getTokenIndex()与源码行号严格对应(后续错误报告依赖此);PRINTF单独作为关键字,而非IDENTIFIER,这样在 visitor 中可直接判断是否为标准库调用,无需查符号表;STRING_LITERAL不支持\n\t等转义,教学场景中学生写的字符串基本是"hello",强行支持\会引入额外 lexer 状态机,得不偿失。
2.3 用 Visitor 模式构建 AST:为什么不用 Listener?因为需要精确控制遍历顺序和返回值
ANTLR 默认生成Listener(事件驱动,无返回值)和Visitor(访问者模式,每个 visit 方法可返回任意类型)。对于编译器,我们必须用Visitor——因为:
visitExpression()必须返回一个ExprNode对象,供父节点(如visitIfStatement())组装条件表达式树;visitFunctionDefinition()必须返回一个FunctionNode,其中包含参数列表、局部变量符号表、IR 指令列表;- 在类型检查阶段,
visitBinaryExpression()需要返回Type枚举(INT_TYPE,CHAR_TYPE),用于检查int + char是否合法。
我们定义核心 AST 节点基类(精简版):
// ASTNode.java public abstract class ASTNode { public final int line; public final int column; public ASTNode(int line, int column) { this.line = line; this.column = column; } } // BinaryOpNode.java public class BinaryOpNode extends ASTNode { public final BinaryOp op; // enum: PLUS, MINUS, MUL, DIV, EQ, NEQ, ... public final ExprNode left; public final ExprNode right; public BinaryOpNode(int line, int column, BinaryOp op, ExprNode left, ExprNode right) { super(line, column); this.op = op; this.left = left; this.right = right; } } // FunctionNode.java public class FunctionNode extends ASTNode { public final String name; public final Type returnType; public final List<ParamNode> params; public final List<VarDeclNode> locals; public final List<StmtNode> body; public FunctionNode(int line, int column, String name, Type returnType, List<ParamNode> params, List<VarDeclNode> locals, List<StmtNode> body) { super(line, column); this.name = name; this.returnType = returnType; this.params = params; this.locals = locals; this.body = body; } }对应的C89BaseVisitor子类关键方法:
// C89ToASTVisitor.java public class C89ToASTVisitor extends C89BaseVisitor<ASTNode> { @Override public ASTNode visitFunctionDefinition(C89Parser.FunctionDefinitionContext ctx) { String funcName = ctx.declarator().IDENTIFIER().getText(); Type returnType = visit(ctx.typeSpecifier()) instanceof IntType ? Type.INT : Type.CHAR; // 解析参数列表(此处简化,实际需遍历 ctx.parameterTypeList()) List<ParamNode> params = new ArrayList<>(); if (ctx.parameterTypeList() != null && ctx.parameterTypeList().parameterList() != null) { for (C89Parser.ParameterListContext p : ctx.parameterTypeList().parameterList()) { // ... 提取参数名和类型 } } // 解析函数体 List<StmtNode> body = new ArrayList<>(); for (C89Parser.StatementContext stmtCtx : ctx.compoundStatement().statement()) { StmtNode stmt = (StmtNode) visit(stmtCtx); body.add(stmt); } return new FunctionNode( ctx.getStart().getLine(), ctx.getStart().getCharPositionInLine(), funcName, returnType, params, new ArrayList<>(), // 局部变量暂空,后续语义分析填充 body ); } @Override public ASTNode visitBinaryExpression(C89Parser.BinaryExpressionContext ctx) { ExprNode left = (ExprNode) visit(ctx.left); ExprNode right = (ExprNode) visit(ctx.right); BinaryOp op = parseBinaryOp(ctx.op.getText()); // 辅助方法映射字符串到 enum return new BinaryOpNode( ctx.getStart().getLine(), ctx.getStart().getCharPositionInLine(), op, left, right ); } // 其他 visit 方法略... }关键逻辑说明:
- 每个
visitXxx()方法都从ctx.getStart()获取line和column,这是错误定位的唯一可靠来源(别信ctx.getText().length()计算位置,ANSI 转义、UTF-8 多字节都会崩); visitFunctionDefinition()中,ctx.declarator().IDENTIFIER()是安全的——因为我们在语法中定义了declarator : IDENTIFIER,ANTLR 保证其存在;BinaryOpNode的构造参数left/right是ExprNode,而非ASTNode,这是类型安全的体现:只有表达式上下文才能产生表达式节点,避免if (x) { y = 1; }中把赋值语句误当表达式。
3. 语义分析与符号表:用 HashMap 实现作用域链,搞定变量声明/使用匹配与类型检查
3.1 符号表设计:为什么不用 ConcurrentHashMap?因为单线程编译器不需要
编译过程是严格的单线程流水线:词法 → 语法 → 语义 → IR 生成 → 汇编输出。并发写符号表不仅没收益,还会因锁竞争拖慢速度。我们用最朴素的HashMap<String, Symbol>,配合作用域嵌套(Scope)实现:
// Scope.java public class Scope { private final Scope parent; // 外层作用域,null 表示全局 private final Map<String, Symbol> symbols; // 当前作用域的符号 public Scope(Scope parent) { this.parent = parent; this.symbols = new HashMap<>(); } // 在当前作用域插入符号(变量、函数) public void define(String name, Symbol symbol) { symbols.put(name, symbol); } // 查找符号:先查当前,再查父作用域,直到全局 public Symbol resolve(String name) { Symbol sym = symbols.get(name); if (sym != null) return sym; if (parent != null) return parent.resolve(name); return null; // 未声明 } public boolean isGlobal() { return parent == null; } }Symbol类封装变量/函数的元信息:
// Symbol.java public class Symbol { public final String name; public final Type type; public final boolean isFunction; public final int offset; // 栈偏移(局部变量)或地址(全局变量) public final List<ParamNode> params; // 仅函数有 public Symbol(String name, Type type, boolean isFunction, int offset) { this(name, type, isFunction, offset, Collections.emptyList()); } public Symbol(String name, Type type, boolean isFunction, int offset, List<ParamNode> params) { this.name = name; this.type = type; this.isFunction = isFunction; this.offset = offset; this.params = params; } }3.2 语义分析 Visitor:在遍历 AST 时同步构建符号表并报错
我们继承C89BaseVisitor<Void>(Void 表示不返回值,只做副作用),在visitFunctionDefinition()中创建新作用域,在visitVarDecl()中插入符号,在visitIdentifierExpr()中查找符号:
// SemanticAnalyzer.java public class SemanticAnalyzer extends C89BaseVisitor<Void> { private Scope currentScope; private final List<Diagnostic> errors; // 错误收集器 public SemanticAnalyzer() { this.currentScope = new Scope(null); // 全局作用域 this.errors = new ArrayList<>(); } @Override public Void visitFunctionDefinition(C89Parser.FunctionDefinitionContext ctx) { String funcName = ctx.declarator().IDENTIFIER().getText(); // 检查函数是否已定义(重复定义) if (currentScope.resolve(funcName) != null) { errors.add(new Diagnostic( Diagnostic.Level.ERROR, ctx.getStart().getLine(), ctx.getStart().getCharPositionInLine(), "redefinition of function '" + funcName + "'" )); return null; } // 创建函数符号(暂不填参数,后面解析) Symbol funcSym = new Symbol(funcName, Type.INT, true, 0); currentScope.define(funcName, funcSym); // 进入函数作用域(新建 Scope,父为 currentScope) Scope funcScope = new Scope(currentScope); this.currentScope = funcScope; // 遍历参数列表,插入参数符号 if (ctx.parameterTypeList() != null) { for (C89Parser.ParameterListContext p : ctx.parameterTypeList().parameterList()) { String paramName = p.IDENTIFIER().getText(); Type paramType = p.typeSpecifier().getText().equals("int") ? Type.INT : Type.CHAR; funcScope.define(paramName, new Symbol(paramName, paramType, false, 0)); } } // 遍历函数体,此时 currentScope 是 funcScope visit(ctx.compoundStatement()); // 函数体结束,恢复外层作用域 this.currentScope = currentScope.parent; return null; } @Override public Void visitVarDecl(C89Parser.VarDeclContext ctx) { String varName = ctx.IDENTIFIER().getText(); Type varType = ctx.typeSpecifier().getText().equals("int") ? Type.INT : Type.CHAR; // 检查是否重复声明 if (currentScope.resolve(varName) != null) { errors.add(new Diagnostic( Diagnostic.Level.ERROR, ctx.getStart().getLine(), ctx.getStart().getCharPositionInLine(), "redefinition of variable '" + varName + "'" )); return null; } // 插入符号(局部变量 offset 从 -4 开始,每声明一个减 4) int offset = -4 * (currentScope.symbols.size() + 1); currentScope.define(varName, new Symbol(varName, varType, false, offset)); return null; } @Override public Void visitIdentifierExpr(C89Parser.IdentifierExprContext ctx) { String ident = ctx.IDENTIFIER().getText(); Symbol sym = currentScope.resolve(ident); if (sym == null) { errors.add(new Diagnostic( Diagnostic.Level.ERROR, ctx.getStart().getLine(), ctx.getStart().getCharPositionInLine(), "use of undeclared identifier '" + ident + "'" )); } return null; } // 其他 visit 方法略... }参数说明:
currentScope是当前活跃作用域,visitFunctionDefinition()中先new Scope(currentScope)创建新作用域,再this.currentScope = funcScope切换,最后this.currentScope = currentScope.parent恢复——这是标准的作用域链管理;offset计算用-4 * (size + 1)是 x86-64 栈帧约定(每个int占 4 字节,从%rbp-4开始),实际生成汇编时会用这个 offset 访问局部变量;Diagnostic是自定义错误类,包含line/column/message,后续可格式化为error: line 5:12: use of undeclared identifier 'x'。
3.3 类型检查:为什么int + char合法,而int * char*不合法?
C89 的隐式类型转换规则(Usual Arithmetic Conversions)必须实现,否则int a=1; char b=2; a+b;会报错。我们在visitBinaryExpression()中加入检查:
@Override public Void visitBinaryExpression(C89Parser.BinaryExpressionContext ctx) { ExprNode left = (ExprNode) visit(ctx.left); ExprNode right = (ExprNode) visit(ctx.right); Type leftType = getType(left); Type rightType = getType(right); // 仅对算术运算符做类型提升 if (ctx.op.getType() == C89Lexer.PLUS || ctx.op.getType() == C89Lexer.MINUS || ctx.op.getType() == C89Lexer.MUL || ctx.op.getType() == C89Lexer.DIV) { // C89 规则:char/short 提升为 int,然后左右类型必须相同 Type promotedLeft = promoteType(leftType); Type promotedRight = promoteType(rightType); if (promotedLeft != promotedRight) { errors.add(new Diagnostic( Diagnostic.Level.ERROR, ctx.getStart().getLine(), ctx.getStart().getCharPositionInLine(), "invalid operands to binary " + ctx.op.getText() + " (have '" + leftType + "' and '" + rightType + "')" )); } } return null; } private Type promoteType(Type t) { if (t == Type.CHAR) return Type.INT; return t; }血泪经验:很多初学者以为char是 1 字节所以不能参与运算,其实 C 标准强制提升为int。不实现这个,printf("%d", 'a' + 1);就会直接报错,而这是最基础的教学案例。
4. 生成汇编代码:用 ASM 库生成 x86-64 NASM 格式,绕过链接器直出可执行文件
4.1 为什么选 ASM 库而非手写字符串拼接?因为指令编码、寄存器分配、栈帧管理全是坑
有人觉得“汇编就是字符串”,写个StringBuilder.append("mov eax, 3").append("\n")就完事。但很快你会遇到:
mov rax, 1234567890123456789是 64 位立即数,NASM 要求mov rax, qword 1234567890123456789,漏qword直接报错;lea rax, [rbp-4]的寻址模式,手写容易丢[ ]或写成rbp-4;- 函数调用时,
rdi,rsi,rdx,rcx,r8,r9,r10,r11是 caller-saved,rbx,r12-r15是 callee-saved,不保存/恢复会破坏调用约定; main函数必须以ret结尾,但printf调用后若不add rsp, 8清理栈,程序崩溃。
ASM 库(我们用org.ow2.asm:asm-tree:9.6)把这些全抽象了:
<dependency> <groupId>org.ow2.asm</groupId> <artifactId>asm-tree</artifactId> <version>9.6</version> </dependency>它提供MethodNode、InsnList、VarInsnNode等类,让你像操作 JVM 字节码一样操作 x86 指令(概念映射:VarInsnNode≈mov eax, [rbp-4],MethodInsnNode≈call printf)。
4.2 生成main函数汇编:从 AST FunctionNode 到 NASM 指令流
我们定义CodeGenerator类,继承C89BaseVisitor<Void>,在visitFunctionDefinition()中生成函数入口:
// CodeGenerator.java public class CodeGenerator extends C89BaseVisitor<Void> { private final PrintWriter out; // 输出到 .s 文件 private final Map<String, Integer> localVarOffsets; // 变量名 → 栈偏移 public CodeGenerator(PrintWriter out) { this.out = out; this.localVarOffsets = new HashMap<>(); } @Override public Void visitFunctionDefinition(C89Parser.FunctionDefinitionContext ctx) { String funcName = ctx.declarator().IDENTIFIER().getText(); // 输出函数标签 out.printf("%s:\n", funcName); out.println(" push rbp"); out.println(" mov rbp, rsp"); // 分配栈空间:每个局部变量占 4 字节,按声明顺序从 rbp-4 开始 int localVarCount = countLocalVars(ctx.compoundStatement()); if (localVarCount > 0) { out.printf(" sub rsp, %d\n", localVarCount * 4); } // 遍历函数体生成指令 visit(ctx.compoundStatement()); // 函数返回:pop rbp; ret out.println(" pop rbp"); out.println(" ret"); return null; } @Override public Void visitReturnStatement(C89Parser.ReturnStatementContext ctx) { if (ctx.expression() != null) { // 生成表达式求值,结果在 eax visit(ctx.expression()); // 将 eax 移到返回寄存器(x86-64 用 rax) out.println(" mov rax, eax"); } return null; } @Override public Void visitBinaryExpression(C89Parser.BinaryExpressionContext ctx) { // 先计算右操作数(入栈),再左操作数(入栈),再运算 visit(ctx.right); out.println(" push rax"); visit(ctx.left); out.println(" pop rbx"); switch (ctx.op.getType()) { case C89Lexer.PLUS: out.println(" add eax, ebx"); break; case C89Lexer.MINUS: out.println(" sub eax, ebx"); break; case C89Lexer.MUL: out.println(" imul eax, ebx"); break; case C89Lexer.DIV: out.println(" cdq"); // 符号扩展 edx:eax out.println(" idiv ebx"); break; } return null; } @Override public Void visitPrintfCall(C89Parser.PrintfCallContext ctx) { // printf 第一个参数是字符串地址,需加载到 rdi if (ctx.expression() != null) { visit(ctx.expression()); // 假设字符串字面量已存为全局数据段 out.println(" mov rdi, str_literal_1"); // 简化:硬编码字符串标签 } else { out.println(" mov rdi, str_literal_0"); // 空字符串 } out.println(" call printf"); out.println(" add rsp, 8"); // 清理栈(printf 是 cdecl 调用约定) return null; } // 其他 visit 方法略... }关键逻辑说明:
push rbp; mov rbp, rsp是标准 x86-64 栈帧建立,sub rsp, N为局部变量分配空间;visitBinaryExpression()中,先visit(right)再visit(left),是因为 C 表达式求值顺序是未定义的,但为简单起见,我们强制右→左,确保a-b中b先入栈;printf调用后add rsp, 8是必须的——因为call printf会把返回地址压栈(8 字节),而printf本身不清理参数栈(cdecl 约定),调用者负责;- 字符串字面量
str_literal_1需在.data段定义,这部分由DataSectionGenerator类单独生成。
4.3 生成.data段:把STRING_LITERAL提取为全局只读数据
// DataSectionGenerator.java public class DataSectionGenerator extends C89BaseVisitor<Void> { private final PrintWriter out; private int stringCounter = 0; public DataSectionGenerator(PrintWriter out) { this.out = out; } @Override public Void visitStringLiteral(C89Parser.StringLiteralContext ctx) { String content = ctx.STRING_LITERAL().getText(); // 去掉首尾双引号,转义 \" 为 " String unquoted = content.substring(1, content.length() - 1).replace("\\\"", "\""); String label = "str_literal_" + (stringCounter++); out.printf("%s: db \"%s\", 0\n", label, unquoted); return null; } }最终生成的.s文件结构:
section .data str_literal_0: db "hello", 0 section .text global main extern printf main: push rbp mov rbp, rsp mov rdi, str_literal_0 call printf add rsp, 8 pop rbp ret用nasm -f elf64 hello.s && gcc -o hello hello.o即可生成可执行文件。这就是“基于 Java 的 C 编译器”的最小可行闭环:Java 代码读入hello.c,输出hello.s,系统工具链完成剩余工作。
5. 避坑指南:编译器开发中 5 个高频翻车现场与后悔药
5.1 现象:ANTLR 报错no viable alternative at input 'int main',但语法文件明明写了functionDefinition
原因:INT和IDENTIFIER的 lexer 规则顺序错了。如果IDENTIFIER定义在INT之前,int会被识别为IDENTIFIER而非INT关键字,导致int main(){}无法匹配functionDefinition(它要求typeSpecifier是INT)。
解决:在.g4文件中,所有关键字规则必须放在IDENTIFIER规则之前。ANTLR 按规则出现顺序匹配,第一个匹配的规则胜出。
5.2 现象:printf("hello")编译后输出乱码或段错误
原因:字符串字面量未正确存入.data段,或mov rdi, str_literal_0中的str_literal_0标签名与.data段定义不一致(大小写、下划线)。更隐蔽的是:printf要求字符串以\0结尾,但STRING_LITERAL规则未自动添加\0。
解决:
- 在
DataSectionGenerator.visitStringLiteral()中,db "%s", 0的0必须显式写出; - 标签名统一用
str_literal_N,生成.s时确保.data段和.text段引用完全一致; - 用
objdump -s hello.o检查.data段内容是否包含"hello\0"。
5.3 现象:int a=1,b=2; a+b;编译通过,但int a=1; char b=2; a+b;报类型错误
原因:未实现 C89 的“整型提升”(Integer Promotion)规则。char在运算前必须提升为int,否则a+b的左右操作数类型不同(intvschar),类型检查失败。
解决:在SemanticAnalyzer.visitBinaryExpression()中,对char/short类型调用promoteType(),统一转为INT后再比较。
5.4 现象:if (x) { y = 1; } else { y = 2; }中y在else分支被使用时报“undeclared identifier”
原因:作用域管理错误。if和else的compoundStatement应在同一个作用域内,但你的visitIfStatement()可能为每个分支新建了Scope。C 语言中if的{}是一个作用域,else的{}是另一个独立作用域,y在if分支声明后,在else分支不可见。
解决:visitIfStatement()不应为if/else创建新作用域;变量声明必须在if外层作用域(即if语句所在的作用域)中完成。教学项目中,建议禁止在if/while内声明变量,强制写成:
int y; if (x) y = 1; else y = 2;5.5 现象:编译大文件(>10KB)时 JVM 报OutOfMemoryError: Java heap space
原因:ANTLR 的CommonTokenStream和 AST 节点在内存中保留全部 token 和节点引用,大文件导致对象过多。这不是代码 bug,是 Java 内存模型限制。
解决:
- 启
本文还有配套的精品资源,点击获取