Commit 31911f51 authored by Data Governance Dev's avatar Data Governance Dev

feat(data-dict): Phase 1 - 单文件脚本完成 smart-build 数据字典与首轮分析

第一阶段,针对 smart-build 数据库做一次性手工治理:
- fetch_data_dictionary.py:拉取 information_schema,输出 data_dictionary.json
- check_empty_fields.py:逐表扫 NULL/空字符串比例,输出高空字段报告
- verify_merge_redundancy.py + .sql:连库验证合并候选的实际行数
- validate_standard_fields.py:对身份证/手机/统一社会信用代码等抽样校验
- check_field_length.py:识别 VARCHAR 超出固定长度的字段(基于国标)
- check_uncommented_fields.py:扫描无注释字段,按字段名启发式推测
- generate_report.py:把以上 JSON 汇总成 Word 报告

Why:先把\"用什么工具看什么数据\"跑通,再抽象成可复用工作流。
How:每个脚本独立运行,输出落 JSON 到 data_dictionary/,最终由
     generate_report.py 汇总成 数据治理报告_smart-build.docx。
parent 20dc9503
# 数据字典获取工具
## 用途
连接 `smart-build` 数据库(MySQL),自动提取全部表和字段的元数据信息,输出为 JSON 文件供后续数据治理分析使用。
## 依赖
- Python 3.7+
- pymysql
```bash
pip install pymysql
```
## 使用方法
```bash
cd data_dictionary
python fetch_data_dictionary.py
```
执行时会提示输入数据库密码(输入不显示,按 Enter 确认)。
## 输出文件
执行成功后在当前目录生成两个 JSON 文件:
| 文件 | 内容 |
|------|------|
| `data_dictionary.json` | 所有表的所有字段详情 |
| `table_summary.json` | 表级别汇总信息 |
### data_dictionary.json 字段说明
| 字段 | 含义 |
|------|------|
| `table_name` | 表名 |
| `column_name` | 字段名 |
| `ordinal_position` | 字段序号(在表中的位置) |
| `column_type` | 完整列类型(如 `varchar(50)`) |
| `data_type` | 数据类型(如 `varchar`) |
| `char_max_length` | 字符最大长度 |
| `numeric_precision` | 数值精度 |
| `numeric_scale` | 小数位数 |
| `is_nullable` | 是否可为空(YES/NO) |
| `column_default` | 默认值 |
| `column_comment` | 字段注释 |
| `extra` | 额外信息(如 `auto_increment`) |
| `table_comment` | 所属表的注释 |
### table_summary.json 字段说明
| 字段 | 含义 |
|------|------|
| `TABLE_NAME` | 表名 |
| `TABLE_TYPE` | 表类型(BASE TABLE / VIEW) |
| `ENGINE` | 存储引擎 |
| `TABLE_ROWS` | 估计行数 |
| `DATA_LENGTH` | 数据大小(字节) |
| `INDEX_LENGTH` | 索引大小(字节) |
| `TABLE_COMMENT` | 表注释 |
| `CREATE_TIME` | 创建时间 |
| `UPDATE_TIME` | 最后更新时间 |
## 数据库连接信息
| 项目 | 值 |
|------|-----|
| 主机 | 192.168.20.10 |
| 端口 | 3306 |
| 用户 | root |
| 数据库 | smart-build |
> 连接信息硬编码在脚本的 `DB_CONFIG` 字典中,如需修改请编辑脚本。
---
## 操作记录
### Step 1: 获取数据字典 ✅ (2026-08-03)
执行 `fetch_data_dictionary.py`,获取 smart-build 库全部 149 张表、3148 个字段的元数据。
### Step 2: 表合并与冗余字段分析 ✅ (2026-08-03)
基于数据字典分析发现:
**合并候选(3组):**
| ID | 表 | 相似度 | 建议 |
|----|-----|--------|------|
| MERGE-001 | `em_gas` + `em_water_supply` | 100% | 24字段完全一致,合并为 `em_infrastructure` 表,加 `type` 字段区分 |
| MERGE-002 | `t_project_*_record` ×5 | 100% | 31字段完全一致,合并为 `t_project_business_record`,加 `record_type` 字段区分 |
| MERGE-003 | `biz_history_field_label` + `biz_history_focus_field` | 83% | 高度相似,考虑合并 |
**废弃候选:**
- **强候选**:19张旧版 `t_*` 人才系统表(全部行数=0,新版 `t_talent_*` 已替代)
- **弱候选**:12张行数为0且无替代的表(mall、project、ai 等模块)
**冗余字段(Top 5):**
| 字段 | 出现表数 | 风险 |
|------|----------|------|
| `project_name` | 30 | 高 — 应通过 project_id 关联 |
| `enterprise_name` | 13 | 高 — 应通过 enterprise_id 关联 |
| `site_code` + `site_name` | 10 | 中 — 成对冗余 |
| `xzqhbm` | 20 | 中 — 已有 c_bri_xzqh 字典表 |
| `user_id` / `talent_id` | 14 / 15 | 中 — 两套ID体系并存 |
详细分析结果见 `findings_table_merge_redundancy.json`。
### Step 3: 数据验证 ✅ (2026-08-03)
执行 `verify_merge_redundancy.py`,11项检查全部通过。原始数据见 `verify_results.json`,最终结论见 `findings_verified.json`。
**确认合并(3组):**
| ID | 表 | 数据量 | 结论 |
|----|-----|--------|------|
| MERGE-001 | `em_gas` + `em_water_supply` | 8行 + 7行 | ✅ 确认合并,加 `infrastructure_type` 字段 |
| MERGE-002 | `t_project_*_record` ×5 | 合计7行 | ✅ 确认合并,加 `record_type` 字段 |
| MERGE-003 | `biz_history_field_label` + `biz_history_focus_field` | 38行 + 13行 | ✅ 建议评估后合并 |
**确认废弃:**
- **强候选**:19张旧版 `t_*` 表全部0行 ✅,新版 `t_talent_*` 正常运行
- **弱候选**:11张0行表(`t_ai_llm_provider` 实际有1行,排除)
**确认冗余字段(5个):**
| 字段 | 表数 | 关键发现 |
|------|------|----------|
| `project_name` | 30 | 13张表同时存 project_id+project_name,冗余维护 |
| `enterprise_name` | 13 | `t_talent_archive`(37155行) 填充率仅1.7% |
| `site_code/site_name` | 10 | **数据已不一致**:同code不同name |
| `xzqhbm` | 20 | 发现孤儿编码;字符集不匹配致部分表无法JOIN |
| `user_id/talent_id` | 14/15 | 两套ID并存;`t_talent_evaluation`/`t_talent_offer` 的 talent_id 填充率0% |
**数据质量问题(3个):**
| ID | 问题 | 严重程度 |
|----|------|----------|
| DQ-001 | em_*/ex_* 与 c_bri_xzqh 字符集不一致(unicode_ci vs general_ci) | 中 |
| DQ-002 | `c_bri_geo_boundary` 20字段 + 表注释全部为空(1804行数据) | 高 |
| DQ-003 | INFORMATION_SCHEMA.TABLE_ROWS 估算不准(0行实际有数据) | 低 |
**下一步:** 进入 Step 4 — 字段规范检查(身份证号、手机号、统一社会信用代码、地区编码等)。
This diff is collapsed.
This diff is collapsed.
This diff is collapsed.
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""
数据字典获取脚本
功能:连接 MySQL 数据库,读取 smart-build 库中所有表的字段信息,
并以 JSON 格式保存到本地文件。
输出字段:
- 表名 (TABLE_NAME)
- 字段名 (COLUMN_NAME)
- 字段注释 (COLUMN_COMMENT)
- 数据类型 (COLUMN_TYPE / DATA_TYPE)
- 字符最大长度 (CHARACTER_MAXIMUM_LENGTH)
- 是否可为空 (IS_NULLABLE)
- 默认值 (COLUMN_DEFAULT)
- 字段序号 (ORDINAL_POSITION)
- 表注释 (TABLE_COMMENT)
"""
import json
import os
import sys
from datetime import datetime
from getpass import getpass
try:
import pymysql
except ImportError:
print("错误:缺少 pymysql 库,请执行: pip install pymysql")
sys.exit(1)
# ── 配置 ──────────────────────────────────────────────
DB_CONFIG = {
"host": "192.168.20.10",
"port": 3306,
"user": "root",
"database": "smart-build",
"charset": "utf8mb4",
}
OUTPUT_FILE = "data_dictionary.json"
OUTPUT_FILE_TABLES = "table_summary.json"
# ── 获取密码 ──────────────────────────────────────────
def get_db_password():
"""以安全方式获取数据库密码(不回显)"""
print("=" * 50)
print(f" 连接目标: {DB_CONFIG['user']}@{DB_CONFIG['host']}:{DB_CONFIG['port']}")
print(f" 数据库: {DB_CONFIG['database']}")
print("=" * 50)
password = getpass("请输入数据库密码(输入不显示): ")
if not password:
print("错误:密码不能为空")
sys.exit(1)
return password
# ── 连接数据库 ────────────────────────────────────────
def connect_db(password):
"""建立数据库连接"""
try:
conn = pymysql.connect(
host=DB_CONFIG["host"],
port=DB_CONFIG["port"],
user=DB_CONFIG["user"],
password=password,
database=DB_CONFIG["database"],
charset=DB_CONFIG["charset"],
connect_timeout=10,
)
print("✓ 数据库连接成功\n")
return conn
except pymysql.err.OperationalError as e:
print(f"错误:无法连接数据库 — {e}")
sys.exit(1)
# ── 查询字段详情 ──────────────────────────────────────
def fetch_column_details(cursor):
"""
从 INFORMATION_SCHEMA.COLUMNS 获取当前库所有字段信息。
返回: list[dict]
"""
sql = """
SELECT
c.TABLE_NAME AS table_name,
c.COLUMN_NAME AS column_name,
c.ORDINAL_POSITION AS ordinal_position,
c.COLUMN_TYPE AS column_type,
c.DATA_TYPE AS data_type,
c.CHARACTER_MAXIMUM_LENGTH AS char_max_length,
c.NUMERIC_PRECISION AS numeric_precision,
c.NUMERIC_SCALE AS numeric_scale,
c.IS_NULLABLE AS is_nullable,
c.COLUMN_DEFAULT AS column_default,
c.COLUMN_COMMENT AS column_comment,
c.EXTRA AS extra,
t.TABLE_COMMENT AS table_comment
FROM
INFORMATION_SCHEMA.COLUMNS AS c
LEFT JOIN
INFORMATION_SCHEMA.TABLES AS t
ON c.TABLE_SCHEMA = t.TABLE_SCHEMA
AND c.TABLE_NAME = t.TABLE_NAME
WHERE
c.TABLE_SCHEMA = %s
ORDER BY
c.TABLE_NAME,
c.ORDINAL_POSITION
"""
cursor.execute(sql, (DB_CONFIG["database"],))
rows = cursor.fetchall()
columns_meta = [desc[0] for desc in cursor.description]
result = []
for row in rows:
record = dict(zip(columns_meta, row))
# 类型转换,确保 JSON 可序列化
for key in ("char_max_length", "numeric_precision", "numeric_scale"):
if record.get(key) is not None:
record[key] = int(record[key])
if record.get("ordinal_position") is not None:
record["ordinal_position"] = int(record["ordinal_position"])
result.append(record)
return result
# ── 查询表汇总信息 ────────────────────────────────────
def fetch_table_summary(cursor):
"""
从 INFORMATION_SCHEMA.TABLES 获取表级别汇总信息。
返回: list[dict]
"""
sql = """
SELECT
TABLE_NAME,
TABLE_TYPE,
ENGINE,
TABLE_ROWS,
DATA_LENGTH,
INDEX_LENGTH,
TABLE_COMMENT,
CREATE_TIME,
UPDATE_TIME
FROM
INFORMATION_SCHEMA.TABLES
WHERE
TABLE_SCHEMA = %s
ORDER BY
TABLE_NAME
"""
cursor.execute(sql, (DB_CONFIG["database"],))
rows = cursor.fetchall()
columns_meta = [desc[0] for desc in cursor.description]
result = []
for row in rows:
record = dict(zip(columns_meta, row))
# 时间字段转字符串
for t_field in ("CREATE_TIME", "UPDATE_TIME"):
if record.get(t_field) is not None:
record[t_field] = record[t_field].strftime("%Y-%m-%d %H:%M:%S")
# 数值转 int
for n_field in ("TABLE_ROWS", "DATA_LENGTH", "INDEX_LENGTH"):
if record.get(n_field) is not None:
record[n_field] = int(record[n_field])
result.append(record)
return result
# ── 保存 JSON ─────────────────────────────────────────
def save_json(data, filepath, label):
script_dir = os.path.dirname(os.path.abspath(__file__))
full_path = os.path.join(script_dir, filepath)
with open(full_path, "w", encoding="utf-8") as f:
json.dump(data, f, ensure_ascii=False, indent=2)
print(f"✓ {label}已保存至: {full_path} ({len(data)} 条记录)")
# ── 统计输出 ──────────────────────────────────────────
def print_stats(column_details, table_summary):
"""输出简要统计信息到控制台"""
tables_in_columns = set(r["table_name"] for r in column_details)
print(f"\n{'─' * 40}")
print(f" 表总数(数据字典): {len(tables_in_columns)}")
print(f" 字段总数: {len(column_details)}")
print(f" 表总数(汇总): {len(table_summary)}")
print(f"{'─' * 40}\n")
# ── 主流程 ────────────────────────────────────────────
def main():
password = get_db_password()
conn = connect_db(password)
try:
with conn.cursor() as cursor:
# 1. 获取字段详情
print("正在获取字段详情...")
column_details = fetch_column_details(cursor)
save_json(column_details, OUTPUT_FILE, "字段详情")
# 2. 获取表汇总
print("正在获取表汇总信息...")
table_summary = fetch_table_summary(cursor)
save_json(table_summary, OUTPUT_FILE_TABLES, "表汇总")
# 3. 统计数据
print_stats(column_details, table_summary)
finally:
conn.close()
print("数据库连接已关闭。")
if __name__ == "__main__":
main()
This diff is collapsed.
This diff is collapsed.
This diff is collapsed.
This diff is collapsed.
Markdown is supported
0%
or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment