☰
毕业设计级Python入侵检测系统:可复现、可解释、可答辩
2026/10/6 14:33:07 网站建设 项目流程

简介:本资源是一套基于Python实现的网络入侵检测与防御系统完整毕业设计项目,面向计算机及相关专业本科生,适用于课程设计、期末大作业及毕业设计实战场景。项目代码经导师指导并获99分高分评价,结构清晰、注释完备,小白可直接运行调试,有效解决初学者在网络安全实践中的环境搭建、模型调用与防御逻辑实现等核心难点。压缩包共38个文件,含14个核心Python源码(如app.py、routes.py)、Docker容器化部署文件(Dockerfile、docker-compose.yml)、前端静态资源(HTML/JS/CSS)及配置说明(requirements.txt、README.md、yml/json配置),整体仅97KB,轻量易部署。目前已有63人学习下载,资源提供开箱即用的完整工程结构、模块化功能划分(如检测模块、响应模块、Web管理界面)及配套文档说明,助读者快速理解IDS/IPS基本原理并完成可演示的系统开发。

1. 这不是“写个防火墙”:一个能跑通、能调参、能答辩的毕业设计级入侵检测系统长什么样?

很多同学拿到“基于Python的网络入侵检测与防御系统”这个毕设题目,第一反应是去GitHub搜个snort-python或scapy-ids项目,改改UI、加个按钮就交差——结果答辩时被问“你这个阈值0.85是怎么定的?”“流量特征提取用了几维向量?PCA降维保留了多少方差?”当场卡壳。真相是:真正能落地的毕业设计级IDS,核心不在“检测准”,而在“可解释、可复现、可验证”。它不需要达到企业级IPS的吞吐量,但必须让导师看到:你理解TCP三次握手如何被用于构造SYN Flood特征,知道为什么HTTP User-Agent字段在异常检测中比URL长度更稳定,清楚Scapy抓包后Raw层和IP层时间戳的精度差异。本项目用纯Python(无C扩展依赖)、单机部署、支持真实PCAP文件+实时网卡捕获双模式,所有模型训练/检测逻辑封装为可调试函数,文档覆盖从Wireshark抓包到混淆矩阵可视化全流程。适合计算机/网络安全方向本科生,零深度学习基础也能上手,但留足了进阶空间——比如把规则引擎换成轻量级LSTM,或把阈值判定改成动态滑动窗口自适应。


2. 从原始流量到结构化特征:用Scapy+Pandas构建可复现的特征工程流水线

2.1 抓包不是目的,可控采样才是关键:本地网卡监听与PCAP回放的统一接口

毕业设计最常翻车的环节,就是“系统在自己电脑上能跑,换台机器就报错”。根源在于抓包权限、网卡名硬编码、时间戳精度不一致。我们用Scapy抽象出统一入口,屏蔽底层差异:

# feature_extractor.py from scapy.all import sniff, rdpcap, IP, TCP, UDP, Raw import pandas as pd from datetime import datetime def capture_or_load(pcap_path=None, interface="eth0", duration=60, packet_count=1000): """ 统一数据源入口:支持实时抓包 or 加载PCAP - pcap_path: 若非None,直接加载PCAP;否则实时抓包 - interface: Linux下常用eth0/wlan0,Windows下用scapy.interfaces()查 - duration: 实时抓包时长(秒),packet_count: 最大包数(防内存溢出) """ if pcap_path: packets = rdpcap(pcap_path) print(f"[INFO] 加载PCAP: {pcap_path}, 共{len(packets)}个数据包") else: print(f"[INFO] 开始监听网卡 {interface},持续{duration}秒...") packets = sniff(iface=interface, timeout=duration, count=packet_count) print(f"[INFO] 实时捕获完成,共{len(packets)}个数据包") # 过滤掉非IP包(如ARP、LLDP),只保留IPv4 ip_packets = [p for p in packets if IP in p] return ip_packets # 示例:加载自带测试PCAP(避免答辩现场抓不到攻击流量) test_packets = capture_or_load(pcap_path="data/test_attack.pcap")

提示:scapy.interfaces()在Windows需管理员权限运行,Linux下普通用户需sudo setcap cap_net_raw+ep /usr/bin/python3授予权限。毕业答辩演示建议预存PCAP,避免现场网络环境不可控。

2.2 特征提取:为什么不用“包长+端口”这种玄学组合?

初学者常把“每个包的长度、源端口、目的端口”当特征,结果模型AUC只有0.52。问题在于:单包特征无法反映攻击行为的时间局部性。SYN Flood的本质是短时间大量SYN包+极少SYN-ACK响应;SQL注入则表现为HTTP请求中特定payload的周期性出现。我们定义两类特征:

特征类型代表字段计算逻辑为什么选它
会话级特征flow_duration,src_bytes,dst_bytes,syn_count,ack_count基于五元组(src_ip, dst_ip, src_port, dst_port, proto)聚合连续5秒内所有包捕捉TCP连接建立/中断异常,如SYN Flood导致syn_count远高于ack_count
统计级特征pkt_per_sec,bytes_per_sec,dst_port_entropy,ttl_std对每秒内所有包计算速率、端口分布熵、TTL标准差端口熵低说明攻击者扫描固定端口(如SSH爆破),TTL标准差大说明伪造IP来源多样
# feature_engineer.py from collections import defaultdict, Counter import numpy as np def extract_session_features(packets, window_sec=5): """ 提取会话级特征:按五元组+时间窗口聚合 返回DataFrame,每行是一个flow(会话片段) """ flows = defaultdict(list) for pkt in packets: if IP in pkt and (TCP in pkt or UDP in pkt): ip_layer = pkt[IP] transport_layer = pkt[TCP] if TCP in pkt else pkt[UDP] key = ( ip_layer.src, ip_layer.dst, transport_layer.sport, transport_layer.dport, ip_layer.proto ) # 使用pkt.time(浮点秒)做时间分桶 bucket = int(pkt.time // window_sec) * window_sec flows[(key, bucket)].append(pkt) features = [] for (key, bucket), pkt_list in flows.items(): src, dst, sport, dport, proto = key # 计算基础统计 durations = [p.time for p in pkt_list] flow_duration = max(durations) - min(durations) if len(durations) > 1 else 0 src_bytes = sum(len(p) for p in pkt_list if p[IP].src == src) dst_bytes = sum(len(p) for p in pkt_list if p[IP].dst == dst) # TCP标志位计数(关键!) syn_count = sum(1 for p in pkt_list if TCP in p and p[TCP].flags & 0x02) # SYN flag ack_count = sum(1 for p in pkt_list if TCP in p and p[TCP].flags & 0x10) # ACK flag features.append({ 'src_ip': src, 'dst_ip': dst, 'src_port': sport, 'dst_port': dport, 'proto': proto, 'flow_duration': flow_duration, 'src_bytes': src_bytes, 'dst_bytes': dst_bytes, 'syn_count': syn_count, 'ack_count': ack_count, 'pkt_count': len(pkt_list) }) return pd.DataFrame(features) # 示例:对test_packets提取特征 df_features = extract_session_features(test_packets) print(df_features.head())

2.3 特征标准化与标签生成:毕业答辩时最该讲清楚的两件事

特征没标准化,模型权重就失去物理意义;标签没明确定义,答辩时“你这正样本怎么来的”就答不上来。我们采用双标签体系:

  • 粗粒度标签(用于模型训练):label∈ {0: normal, 1: attack},由PCAP文件名隐含(如ddos_synflood.pcap→ label=1)
  • 细粒度标签(用于分析报告):attack_type∈ {syn_flood, port_scan, sql_inject, ...},存于data/labels.csv
# label_generator.py import os import re def generate_labels_from_pcap_dir(pcap_dir): """ 根据PCAP文件名自动打标(毕业设计可接受的合理方案) 规则:文件名含'synflood'→'syn_flood', 'nmap'→'port_scan', 'sqlmap'→'sql_inject' """ label_map = { 'synflood': 'syn_flood', 'nmap': 'port_scan', 'sqlmap': 'sql_inject', 'hydra': 'brute_force', 'normal': 'normal' } labels = [] for f in os.listdir(pcap_dir): if f.endswith('.pcap'): label = 'normal' # 默认正常 for keyword, attack in label_map.items(): if keyword in f.lower(): label = attack break labels.append({'filename': f, 'label': 0 if label=='normal' else 1, 'attack_type': label}) return pd.DataFrame(labels) # 特征标准化:用MinMaxScaler而非StandardScaler,避免负值(如包长不能为负) from sklearn.preprocessing import MinMaxScaler def normalize_features(df, exclude_cols=['src_ip','dst_ip','attack_type']): scaler = MinMaxScaler() numeric_cols = df.select_dtypes(include=[np.number]).columns.tolist() numeric_cols = [c for c in numeric_cols if c not in exclude_cols] df_norm = df.copy() df_norm[numeric_cols] = scaler.fit_transform(df[numeric_cols]) return df_norm, scaler # 示例:生成标签 + 标准化 labels_df = generate_labels_from_pcap_dir("data/") df_norm, scaler = normalize_features(df_features)

3. 检测模型选型:为什么随机森林比XGBoost更适合毕业设计场景?

3.1 毕业设计的三大隐形约束:可解释性、训练速度、依赖精简

XGBoost在Kaggle上AUC常超0.95,但毕业答辩时导师问:“这个feature_importance排序,为什么‘syn_count’排第三而‘pkt_per_sec’排第一?请结合TCP协议栈解释”,你就得现场推导三次握手状态机。而随机森林:

  • 特征重要性直观:syn_count高→SYN Flood可能性大,dst_port_entropy低→端口扫描,结论可直接映射到网络协议;
  • 训练快:1万条flow特征,i5笔记本30秒出模型,不用等答辩前夜还在跑GridSearch;
  • 依赖少:sklearn.ensemble.RandomForestClassifier,无CUDA、无LightGBM编译烦恼。
# model_trainer.py from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import train_test_split from sklearn.metrics import classification_report, confusion_matrix import joblib def train_rf_model(X_train, y_train, n_estimators=100, max_depth=10, random_state=42): """ 训练RF模型:参数已针对小数据集优化 - n_estimators=100: 平衡效果与速度(>200训练慢,<50易过拟合) - max_depth=10: 防止树过深(毕业设计数据量小,深树泛化差) """ model = RandomForestClassifier( n_estimators=n_estimators, max_depth=max_depth, random_state=random_state, n_jobs=-1 # 利用所有CPU核心 ) model.fit(X_train, y_train) return model # 数据准备(假设df_norm已含label列) X = df_norm.drop(['src_ip','dst_ip','label'], axis=1) y = df_norm['label'] X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.3, random_state=42, stratify=y ) rf_model = train_rf_model(X_train, y_train) y_pred = rf_model.predict(X_test) print(classification_report(y_test, y_pred)) # 输出示例: # precision recall f1-score support # 0 0.98 0.97 0.97 892 # 1 0.95 0.96 0.95 308 # accuracy 0.97 1200

3.2 模型可解释性:用SHAP可视化,答辩时直接展示“为什么判为攻击”

毕业设计答辩最加分的环节,不是AUC多高,而是你能指着图说:“看,这个flow的syn_count=128(正常均值是3),贡献了0.73的SHAP值,所以模型判为攻击”。SHAP比sklearn自带feature_importances_更精细:

# explainability.py import shap import matplotlib.pyplot as plt def plot_shap_summary(model, X_test, feature_names, save_path="shap_summary.png"): """ 生成SHAP摘要图:横轴是SHAP值(影响程度),纵轴是特征 红点=高值→正向影响(促发攻击判定),蓝点=低值→负向影响 """ explainer = shap.TreeExplainer(model) shap_values = explainer.shap_values(X_test) plt.figure(figsize=(10, 6)) shap.summary_plot(shap_values[1], X_test, feature_names=feature_names, show=False) plt.title("SHAP Feature Importance (Attack Class)") plt.tight_layout() plt.savefig(save_path, dpi=300, bbox_inches='tight') plt.show() # 调用示例 feature_names = X.columns.tolist() plot_shap_summary(rf_model, X_test, feature_names)

注意:SHAP计算较慢,毕业设计只需在测试集上跑一次生成图,无需实时解释。图中若syn_count和ack_count形成强对比(如syn高+ack低),就是SYN Flood的铁证。


4. 防御动作执行:不是弹窗警告,而是可验证的主动响应链

4.1 防御≠阻断:毕业设计阶段最务实的三类响应动作

企业级IDS对接防火墙API,但毕业设计要的是可验证、可截图、可写进论文的动作。我们定义三级响应:

响应等级动作验证方式代码复杂度
Level 1:日志告警写入alerts.log,含时间、IP、攻击类型、置信度tail -f alerts.log实时查看★☆☆
Level 2:进程终止os.system(f"kill -9 $(lsof -i:{port} -t)")终止可疑端口进程netstat -tuln | grep {port}查无监听★★☆
Level 3:iptables封禁os.system(f"iptables -A INPUT -s {ip} -j DROP")iptables -L INPUT -n | grep {ip}查存在★★★
# defense_executor.py import os import logging from datetime import datetime # 配置日志(Level 1响应) logging.basicConfig( filename='alerts.log', level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s' ) def log_alert(src_ip, attack_type, confidence, threshold=0.7): """Level 1:日志告警""" if confidence >= threshold: msg = f"ALERT: {attack_type} detected from {src_ip} (confidence: {confidence:.3f})" logging.info(msg) print(f"[ALERT] {msg}") def kill_suspicious_process(dst_port, protocol='tcp'): """Level 2:终止可疑进程(需root权限)""" try: # 查找占用dst_port的进程PID cmd = f"lsof -i {protocol}:{dst_port} -t 2>/dev/null" pid = os.popen(cmd).read().strip() if pid: os.system(f"kill -9 {pid}") print(f"[DEFENSE] Killed process {pid} on port {dst_port}") return True except Exception as e: print(f"[ERROR] Failed to kill process: {e}") return False def block_ip_with_iptables(src_ip): """Level 3:iptables封禁(Linux专用)""" try: # 检查是否已封禁 check_cmd = f"iptables -L INPUT -n | grep ' {src_ip} '" if os.system(check_cmd + " >/dev/null 2>&1") != 0: os.system(f"iptables -A INPUT -s {src_ip} -j DROP") print(f"[DEFENSE] Blocked IP {src_ip} via iptables") return True except Exception as e: print(f"[ERROR] iptables block failed: {e}") return False # 示例:当模型预测为攻击且置信度>0.8时触发 if y_pred[0] == 1 and rf_model.predict_proba(X_test.iloc[[0]])[0][1] > 0.8: src_ip = df_norm.iloc[0]['src_ip'] log_alert(src_ip, "syn_flood", 0.85) # kill_suspicious_process(80) # 可选 # block_ip_with_iptables(src_ip) # 需sudo权限

4.2 响应动作的边界控制:为什么毕业设计绝不该自动删库?

曾有学生写os.system("rm -rf /var/www/html/*")当“防御”,答辩时被问“如果误报,删的是教务系统网站怎么办?”——防御动作必须带人工确认开关。我们在配置文件中强制设置:

# config.py class DefenseConfig: ENABLE_LOG_ALERT = True # 日志告警永远开启 ENABLE_PROCESS_KILL = False # 默认关闭,答辩演示时手动设True ENABLE_IPTABLES_BLOCK = False # 默认False,避免误封导师IP # 关键:白名单IP,永不封禁 WHITELIST_IPS = ['192.168.1.100', '127.0.0.1'] # 实验室服务器IP、本机 # 封禁前二次确认(仅开发/演示用) def confirm_block(self, ip): if ip in self.WHITELIST_IPS: print(f"[SKIP] {ip} is in whitelist, skip blocking.") return False confirm = input(f"Confirm block {ip}? (y/N): ") return confirm.lower() == 'y' # 使用示例 cfg = DefenseConfig() if cfg.ENABLE_IPTABLES_BLOCK and cfg.confirm_block("192.168.1.200"): block_ip_with_iptables("192.168.1.200")

血泪经验:答辩前一天务必检查WHITELIST_IPS包含自己电脑IP和实验室服务器IP,否则封禁后连不上演示机。


5. 避坑指南:答辩前夜最常踩的5个坑及救急方案

5.1 现象:Scapy抓包为空列表,len(packets)==0

原因:

  • Windows下未以管理员身份运行Python;
  • Linux下网卡名错误(eth0在新Ubuntu可能是ens33);
  • Wireshark正在同一网卡抓包,导致Scapy抢不到权限。
    解决:
# Linux查网卡名 ip link show | grep "state UP" -A1 # Windows查网卡名(PowerShell) Get-NetAdapter | Where-Object {$_.Status -eq "Up"} | Select-Object Name # 关闭Wireshark再试

5.2 现象:训练时报错ValueError: Input contains NaN

原因:
特征工程中除零(如某flow无包,pkt_per_sec=0/0)导致NaN;
PCAP里有畸形包(如IP头损坏),Scapy解析失败返回None。
解决:

# 在extract_session_features末尾添加清洗 df_features = df_features.replace([np.inf, -np.inf], np.nan) df_features = df_features.dropna() # 删除含NaN行 # 或用0填充(更安全) df_features = df_features.fillna(0)

5.3 现象:模型预测全是0(全判正常)

原因:
标签不平衡(如1000个normal + 10个attack),RF默认class_weight='balanced'未启用;
特征未标准化,syn_count(值1000+)淹没ttl_std(值0.1~2)。
解决:

# 训练时显式设置 model = RandomForestClassifier( class_weight='balanced', # 关键! ... ) # 确保normalize_features已执行

5.4 现象:SHAP图报错'numpy.ndarray' object has no attribute 'shape'

原因:
shap_values返回的是list(二分类时[values_for_class_0, values_for_class_1]),传给summary_plot时需指定用哪一类。
解决:

# 正确用法:取attack class(索引1)的SHAP值 shap.summary_plot(shap_values[1], X_test, ...) # 不是shap_values

5.5 现象:iptables封禁后无法SSH登录

原因:
iptables -A INPUT -s x.x.x.x -j DROP追加规则,但已有ACCEPT规则在前,DROP不生效;
未保存规则,重启后失效。
解决:

# 查看规则顺序 sudo iptables -L INPUT -n --line-numbers # 插入到最前面(确保优先匹配) sudo iptables -I INPUT 1 -s 192.168.1.200 -j DROP # 保存(Ubuntu) sudo iptables-save > /etc/iptables/rules.v4 # 临时清空(救急) sudo iptables -F INPUT

6. 毕业答辩终极技巧:用“三页纸”讲清整个系统的技术纵深

6.1 答辩PPT技术页的黄金结构:问题→解法→证据

别堆代码!导师想看的是你思考的链条。每页只讲1个技术点,按此结构:

区域内容占比示例
左上角(问题)用1句话点出痛点15%“单包特征无法识别SYN Flood的时间局部性”
中部(解法)你的创新/选择理由50%“采用5秒滑动窗口聚合五元组,定义syn_count/ack_count比值作为核心指标”
右下角(证据)一张图或一行命令结果35%SHAP图中syn_count高亮区域 +print(df_features['syn_count'].describe())输出均值/最大值

6.2 让导师眼前一亮的3个细节操作

① 演示时用真实攻击PCAP,但提前标注好“攻击起始时间”
在Wireshark打开ddos_synflood.pcap,标记第12.34秒开始SYN Flood,答辩时直接跳转:“请看这里,我们的系统在12.38秒触发告警,滞后仅0.04秒”。

② 展示特征工程代码时,高亮syn_count计算行并解释协议依据

# TCP flags: 0x02 = SYN, 0x10 = ACK syn_count = sum(1 for p in pkt_list if TCP in p and p[TCP].flags & 0x02)

口头补充:“RFC 793规定SYN包必须设置SYN flag(bit 1),而正常连接中SYN包数应≈ACK包数,比值>5即异常——这比单纯数包长更符合协议本质。”

③ 模型评估不用AUC,用“攻击检出延迟”和“误报率”双指标

# 计算从攻击开始到首次告警的时间差(秒) attack_start = 12.34 first_alert_time = 12.38 detection_latency = first_alert_time - attack_start # 0.04s # 误报率:在1小时正常流量中告警次数 / 总flow数 false_alarm_rate = 3 / 12500 # 0.024%

强调:“企业关注吞吐量,毕业设计关注可验证性——0.04秒检出+0.024%误报,证明特征工程有效。”

我带过7届毕设,最常后悔的是学生花两周调TensorFlow却没搞懂TCP状态机。这个项目真正的价值,不是写出多炫的模型,而是让你亲手把scapy.Packet变成pandas.Series,再变成sklearn.Predictor,最后变成iptables -A INPUT——每一行代码都踩在协议栈的砖石上,而不是浮在API文档的泡沫里。希望帮到你。

本文还有配套的精品资源,点击获取

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询