""" 방법론 자동 모니터링 파이프라인 ================================ 주기적으로 실행하여 신규/개정 방법론을 감지하고 knowledge_base를 자동으로 갱신합니다. 실행: python monitor_methodologies.py # 변경 감지 + 자동 업데이트 python monitor_methodologies.py --check-only # 변경 감지만 (업데이트 없음) python monitor_methodologies.py --scrape # 원격 레지스트리 스크래핑 포함 python monitor_methodologies.py --dry-run # 미리보기 흐름: ① 로컬 방법론 디렉토리 스캔 (CDM/ORS/GS/Verra) ② methodology_state.json 과 비교 → 신규/개정 감지 ③ 변경된 파일 → knowledge_base/ 복사 ④ 변경 있으면 → update_all.py 실행 (ChromaDB 재구축) ⑤ state 파일 갱신 """ from __future__ import annotations import argparse import hashlib import json import shutil import subprocess import sys import time from dataclasses import asdict, dataclass from datetime import datetime from pathlib import Path from typing import Optional # ── 경로 설정 ────────────────────────────────────────────────────────────── ROOT = Path(__file__).parent KNOWLEDGE_BASE = ROOT / "react-agent" / "knowledge_base" STATE_FILE = ROOT / "methodology_state.json" # 레지스트리별 로컬 디렉토리 및 prefix 설정 REGISTRIES = { "CDM_PA": { "dir": ROOT / "cdm_methodologies" / "CDM_PA_Large_Scale_Methodologies", "prefix": "cdm_pa_", "label": "CDM 대규모", }, "CDM_SSC": { "dir": ROOT / "cdm_methodologies" / "CDM_SSC_Small_Scale_Methodologies", "prefix": "cdm_ssc_", "label": "CDM 소규모", }, "CDM_AR": { "dir": ROOT / "cdm_methodologies" / "CDM_AR_SSCAR_Afforestation_Reforestation_Methodologies", "prefix": "cdm_ar_", "label": "CDM 조림/재조림", }, "ORS": { "dir": ROOT / "ors_methodologies", "prefix": "ors_", "label": "ORS (국내 상쇄배출원)", }, "GS": { "dir": ROOT / "gs_methodologies", "prefix": "gs_", "label": "Gold Standard", }, "VERRA": { "dir": ROOT / "verra_methodologies", "prefix": "verra_", "label": "Verra VCS", }, } # knowledge_base에 복사할 파일 확장자 ALLOWED_EXT = {".pdf", ".hwp", ".docx", ".txt"} # 스크래퍼 스크립트 (--scrape 옵션 사용 시) SCRAPERS = { "ORS": ROOT / "scrape_ors.py", "GS": ROOT / "scrape_gs.py", "VERRA": ROOT / "scrape_verra.py", } # ── 데이터 구조 ──────────────────────────────────────────────────────────── @dataclass class FileRecord: registry: str filename: str # 원본 파일명 kb_filename: str # knowledge_base 내 파일명 file_hash: str # MD5 해시 file_size: int last_seen: str # ISO 날짜 in_kb: bool = False # knowledge_base에 복사됐는지 def compute_hash(path: Path) -> str: """파일 MD5 해시 계산""" h = hashlib.md5() with open(path, "rb") as f: for chunk in iter(lambda: f.read(65536), b""): h.update(chunk) return h.hexdigest() def load_state() -> dict[str, FileRecord]: """상태 파일 로드""" if not STATE_FILE.exists(): return {} try: with open(STATE_FILE, "r", encoding="utf-8") as f: raw = json.load(f) return {k: FileRecord(**v) for k, v in raw.items()} except Exception: return {} def save_state(state: dict[str, FileRecord]) -> None: """상태 파일 저장""" with open(STATE_FILE, "w", encoding="utf-8") as f: json.dump({k: asdict(v) for k, v in state.items()}, f, ensure_ascii=False, indent=2) def make_state_key(registry: str, filename: str) -> str: return f"{registry}:{filename}" # ── 스캔 ────────────────────────────────────────────────────────────────── def scan_registry(registry_id: str, config: dict) -> dict[str, FileRecord]: """레지스트리 디렉토리 스캔 → 현재 파일 목록 반환""" dir_path: Path = config["dir"] prefix: str = config["prefix"] records = {} if not dir_path.exists(): return records for file_path in sorted(dir_path.iterdir()): if file_path.suffix.lower() not in ALLOWED_EXT: continue if not file_path.is_file(): continue key = make_state_key(registry_id, file_path.name) file_hash = compute_hash(file_path) kb_filename = prefix + file_path.name records[key] = FileRecord( registry=registry_id, filename=file_path.name, kb_filename=kb_filename, file_hash=file_hash, file_size=file_path.stat().st_size, last_seen=datetime.now().isoformat(), in_kb=(KNOWLEDGE_BASE / kb_filename).exists(), ) return records def scan_all() -> dict[str, FileRecord]: """모든 레지스트리 스캔""" all_records = {} for registry_id, config in REGISTRIES.items(): records = scan_registry(registry_id, config) all_records.update(records) return all_records # ── 변경 감지 ───────────────────────────────────────────────────────────── @dataclass class ChangeReport: new_files: list[FileRecord] # 신규 방법론 updated_files: list[FileRecord] # 개정된 방법론 removed_keys: list[str] # 삭제된 방법론 not_in_kb: list[FileRecord] # 로컬엔 있지만 KB에 없는 파일 def detect_changes( current: dict[str, FileRecord], previous: dict[str, FileRecord], ) -> ChangeReport: """현재 상태 vs 이전 상태 비교""" new_files = [] updated_files = [] not_in_kb = [] for key, rec in current.items(): if key not in previous: new_files.append(rec) elif previous[key].file_hash != rec.file_hash: updated_files.append(rec) elif not rec.in_kb: not_in_kb.append(rec) removed_keys = [k for k in previous if k not in current] return ChangeReport( new_files=new_files, updated_files=updated_files, removed_keys=removed_keys, not_in_kb=not_in_kb, ) # ── Knowledge Base 동기화 ───────────────────────────────────────────────── def sync_to_knowledge_base( records: list[FileRecord], dry_run: bool = False, ) -> list[FileRecord]: """파일들을 knowledge_base/에 복사""" synced = [] KNOWLEDGE_BASE.mkdir(parents=True, exist_ok=True) for rec in records: # 원본 파일 경로 복원 config = REGISTRIES[rec.registry] src = config["dir"] / rec.filename dst = KNOWLEDGE_BASE / rec.kb_filename if not src.exists(): print(f" [경고] 원본 파일 없음: {src}") continue if dry_run: print(f" [DRY-RUN] 복사 예정: {rec.kb_filename}") synced.append(rec) continue shutil.copy2(src, dst) rec.in_kb = True synced.append(rec) print(f" 복사: {rec.kb_filename} ({rec.file_size / 1024:.0f}KB)") return synced def remove_from_knowledge_base( keys: list[str], state: dict[str, FileRecord], dry_run: bool = False, ) -> None: """삭제된 방법론을 knowledge_base에서도 제거""" for key in keys: if key not in state: continue rec = state[key] dst = KNOWLEDGE_BASE / rec.kb_filename if dst.exists(): if dry_run: print(f" [DRY-RUN] 삭제 예정: {rec.kb_filename}") else: dst.unlink() print(f" 삭제: {rec.kb_filename}") # ── 스크래퍼 실행 ──────────────────────────────────────────────────────── def run_scrapers(targets: Optional[list[str]] = None) -> None: """원격 레지스트리 스크래핑 (새 파일 다운로드)""" targets = targets or list(SCRAPERS.keys()) for name in targets: scraper = SCRAPERS.get(name) if not scraper or not scraper.exists(): print(f" [경고] 스크래퍼 없음: {name}") continue print(f"\n 스크래핑: {name}...") result = subprocess.run( [sys.executable, str(scraper)], cwd=str(ROOT), ) if result.returncode != 0: print(f" [경고] {name} 스크래퍼 실패 (종료코드 {result.returncode})") # ── update_all.py 실행 ──────────────────────────────────────────────────── def run_update_pipeline(dry_run: bool = False) -> bool: """update_all.py 실행 (ChromaDB + GraphDB + 배포)""" update_script = ROOT / "update_all.py" if not update_script.exists(): print(" [오류] update_all.py 없음") return False cmd = [sys.executable, str(update_script), "--yes", "--skip-graph", "--skip-community"] if dry_run: cmd.append("--dry-run") print(f"\n update_all.py 실행 중...") result = subprocess.run(cmd, cwd=str(ROOT)) return result.returncode == 0 # ── 리포트 출력 ────────────────────────────────────────────────────────── def print_report(report: ChangeReport, current: dict[str, FileRecord]) -> None: """변경 감지 결과 출력""" total = len(report.new_files) + len(report.updated_files) + len(report.not_in_kb) print("\n" + "=" * 60) print(" 방법론 모니터링 결과") print("=" * 60) # 레지스트리별 통계 registry_counts = {} for rec in current.values(): registry_counts[rec.registry] = registry_counts.get(rec.registry, 0) + 1 print("\n 레지스트리별 보유 방법론:") for rid, config in REGISTRIES.items(): count = registry_counts.get(rid, 0) label = config["label"] print(f" {label:25s}: {count:4d}개") print(f"\n 총 보유 방법론: {len(current)}개") # 변경사항 if report.new_files: print(f"\n [신규] {len(report.new_files)}개 발견:") for rec in report.new_files[:10]: label = REGISTRIES[rec.registry]["label"] print(f" [{label}] {rec.filename[:60]}") if len(report.new_files) > 10: print(f" ... 외 {len(report.new_files)-10}개") if report.updated_files: print(f"\n [개정] {len(report.updated_files)}개 감지:") for rec in report.updated_files[:10]: label = REGISTRIES[rec.registry]["label"] print(f" [{label}] {rec.filename[:60]}") if report.not_in_kb: print(f"\n [미동기] KB에 없는 파일 {len(report.not_in_kb)}개") if report.removed_keys: print(f"\n [삭제] {len(report.removed_keys)}개 방법론 제거됨") if total == 0: print("\n 변경사항 없음 - Knowledge Base 최신 상태") else: print(f"\n → Knowledge Base 갱신 필요: {total}개 파일") print("=" * 60) # ── 메인 ───────────────────────────────────────────────────────────────── def main(): parser = argparse.ArgumentParser( description="방법론 자동 모니터링 파이프라인" ) parser.add_argument("--check-only", action="store_true", help="변경 감지만 (KB 갱신 없음)") parser.add_argument("--scrape", action="store_true", help="원격 레지스트리 스크래핑 후 감지") parser.add_argument("--scrape-only", nargs="+", metavar="REGISTRY", help=f"특정 레지스트리만 스크래핑 ({', '.join(SCRAPERS.keys())})") parser.add_argument("--dry-run", action="store_true", help="미리보기 (실제 파일 변경 없음)") parser.add_argument("--skip-update", action="store_true", help="update_all.py 실행 건너뜀") args = parser.parse_args() print("=" * 60) print(" 방법론 자동 모니터링 파이프라인") print(f" 실행 시각: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}") print("=" * 60) # Step 0: 스크래핑 (옵션) if args.scrape or args.scrape_only: print("\n[Step 0] 원격 레지스트리 스크래핑") targets = args.scrape_only or None run_scrapers(targets) # Step 1: 현재 상태 스캔 print("\n[Step 1] 로컬 방법론 디렉토리 스캔 중...") t0 = time.perf_counter() current_state = scan_all() print(f" 완료: {len(current_state)}개 파일 ({time.perf_counter()-t0:.1f}초)") # Step 2: 이전 상태 로드 & 비교 print("\n[Step 2] 변경사항 감지 중...") previous_state = load_state() report = detect_changes(current_state, previous_state) # 리포트 출력 print_report(report, current_state) # --check-only 이면 여기서 종료 if args.check_only: save_state(current_state) return # Step 3: Knowledge Base 동기화 files_to_sync = report.new_files + report.updated_files + report.not_in_kb if files_to_sync: print(f"\n[Step 3] Knowledge Base 동기화 ({len(files_to_sync)}개)") synced = sync_to_knowledge_base(files_to_sync, dry_run=args.dry_run) # 삭제된 방법론 KB에서도 제거 if report.removed_keys: remove_from_knowledge_base( report.removed_keys, previous_state, dry_run=args.dry_run ) # 상태 파일 갱신 if not args.dry_run: # 현재 상태로 업데이트 (removed는 제외) new_state = {k: v for k, v in current_state.items()} save_state(new_state) print(f"\n 상태 파일 저장: {STATE_FILE}") # Step 4: update_all.py 실행 if not args.skip_update and synced: print(f"\n[Step 4] Knowledge Base 재구축 (update_all.py)") if args.dry_run: print(" [DRY-RUN] update_all.py 실행 건너뜀") else: success = run_update_pipeline(dry_run=False) if success: print(" Knowledge Base 재구축 완료") print("\n 다음 단계:") print(" git push origin main # HuggingFace Spaces 배포") else: print(" [경고] update_all.py 실패 - 수동 확인 필요") else: # 변경 없어도 상태는 저장 (last_seen 갱신) if not args.dry_run: save_state(current_state) print("\n Knowledge Base 갱신 불필요") print(f"\n완료: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\n") if __name__ == "__main__": main()