[자료집] 2026 제3차 AI 안전 정책 연구 세미나 「LLM 안전성 벤치마크의 한국 맥락화」
-
분류연구
-
등록일2026-07-27 17:31:24
-
작성자운영자
-
조회수200
-
인공지능안전연구소는 AI 위험과 AI 안전 분야의 주요 이슈를 국내외 전문가와 함께 논의하고 중장기 연구 아젠다를 발굴하기 위하여 「2026 AI 안전 정책 연구 세미나 시리즈」를 운영하고 있습니다.
그 세 번째 결과로 제3차 세미나 「LLM 안전성 벤치마크의 한국 맥락화(Korean-Contextualized LLM Safety Benchmarks)」의 자료집을 공개합니다.
세미나 개요
주제 LLM 안전성 벤치마크의 한국 맥락화 (Korean-Contextualized LLM Safety Benchmarks)
주최 인공지능안전연구소
일시 2026년 7월 6일(월) 10:00~12:30
장소 인공지능안전연구소 (온·오프라인 동시 진행)
제3차 세미나에서는 한국어와 문화적 맥락을 반영한 LLM 안전성 평가와 벤치마크 개발을 주제로 두 건의 발표가 진행되었습니다.
첫 번째 발표에서 Scale AI의 Mike Lee는 「ROK-FORTRESS: Why Translation-Only Safety Evaluations Fall Short for Korean Deployment」를 통해 한국어와 한국 문화 맥락을 반영한 AI 안전성 평가 벤치마크(ROK-FORTRESS)를 제안하였습니다. 언어와 문화를 독립적으로 통제하는 Transcreation Matrix를 구축하여 1,235개 국가안보·공공안전(NSPS) 문항을 개발하고 14종의 LLM을 평가하였으며, 한국어 입력에서 유해 응답이 감소하는 경향과 언어 효과가 문화 효과보다 약 2.5배 크게 나타나는 결과를 제시하였습니다.
두 번째 발표에서 DATUMO의 Junghoon Kim은 「From Korean Context to Multimodal Safety: Transcreating Culture-Aware LLM Benchmarks」를 통해 한국형 AI 안전성 평가를 위한 Transcreation 기반 벤치마크 구축 방법론(CAGE)과 멀티모달 안전성 평가 방향을 제안하였습니다. 국가별 법·문화·제도·위협 환경에 따라 AI 안전성이 달라져 번역만으로는 충분하지 않다는 점을 지적하고, 문화적 맥락을 재구성하는 CAGE(Culturally Adaptive Red-Teaming Benchmark Generation) 프레임워크와 이미지와 텍스트를 함께 고려하는 멀티모달 Transcreation 파이프라인을 제시하였습니다.
자료집 전문은 첨부파일에서 확인하실 수 있습니다.
The Korea AI Safety Institute (Korea AISI) runs the 2026 AI Safety Policy Research Seminar Series to discuss key issues in AI risk and AI safety with domestic and international experts and to identify medium- and long-term research agendas. As the third installment, we are pleased to share the proceedings of the third seminar, "Korean-Contextualized LLM Safety Benchmarks."
The third seminar was held on July 6, 2026, at the Korea AI Safety Institute, both online and in person, featuring two presentations on AI safety evaluation and benchmarks that reflect Korean language and cultural contexts.
In the first presentation, "ROK-FORTRESS: Why Translation-Only Safety Evaluations Fall Short for Korean Deployment," Mike Lee (Scale AI) proposed the ROK-FORTRESS benchmark for evaluating AI safety in Korean language and cultural contexts. Using a Transcreation Matrix that independently controls for language and culture, he developed 1,235 National Security and Public Safety tasks and evaluated 14 LLMs, finding that Korean prompts consistently reduced harmful outputs and that language effects were about 2.5 times larger than culture effects.
In the second presentation, "From Korean Context to Multimodal Safety: Transcreating Culture-Aware LLM Benchmarks," Junghoon Kim (DATUMO Inc.) proposed a transcreation-based benchmark development methodology (CAGE) and a multimodal AI safety evaluation framework. Arguing that translation-only evaluation is insufficient because AI safety depends on country-specific legal, cultural, institutional, and threat contexts, he introduced the CAGE (Culturally Adaptive Red-Teaming Benchmark Generation) framework and a multimodal transcreation pipeline incorporating both text and images.
The full proceedings are available in the attached file. -
첨부파일