새 모델 공개, 벤치마크 결과, 학습 기법과 논문 등 모델 자체에 관한 소식. New model releases, benchmark results, training techniques, and research papers.
-
How AI is expanding what people do at work
오픈AI가 미국 챗GPT 사용자의 업무 관련 메시지 80만 건 이상을 분석한 결과, 업무 관련 메시지의 16.8%, 특정 직업에 특화된 메시지의 43.5%가 원래 자신의 직무가 아닌 다른 직업의 업무를 처리하는 데 쓰인 것으로 나타났다. 소상공인이 직접 광고 문구를 쓰거나 계약서를 검토하고, 영업직이 고객 데이터를 분석하고, 마케터가 개발자 없이 웹사이트 문제를 고치는 사례가 대표적이다. OpenAI analyzed more than 800,000 U.S. ChatGPT work-related messages and found that 16.8% of work-related messages — and 43.5% of messages tied to a specific occupation — involved tasks normally associated with a different job, a pattern the company calls 'task crossover.' Examples include a small-business owner drafting ad copy or reviewing contracts, a salesperson digging into customer data, and a marketer fixing a website without a developer.
-
Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows
앤스로픽이 7월 24일 새 모델 클로드 오퍼스 5를 출시했다. 입력 토큰당 5달러, 출력 토큰당 25달러로 이전 모델 오퍼스 4.8과 같은 가격을 유지하면서, 에이전트 코딩 벤치마크인 Frontier-Bench v0.1에서 43.3%를 기록해 오퍼스 4.8(18.7%)의 두 배 이상 점수를 냈다. 오퍼스 5는 클로드 맥스의 기본 모델이자 클로드 프로에서 쓸 수 있는 가장 강력한 모델이 됐다. Anthropic released Claude Opus 5 on July 24, keeping pricing at $5 per million input tokens and $25 per million output tokens — unchanged from Opus 4.8 — while more than doubling its score on the Frontier-Bench v0.1 agentic coding benchmark (43.3% vs. 18.7%). Opus 5 is now the default model on Claude Max and the strongest model available on Claude Pro.
-
Making sense of the panic over Chinese AI
중국 문샷AI의 신모델 '키미(Kimi)' 출시가 실리콘밸리와 월가에 중국 AI 경쟁력에 대한 불안을 촉발했다. 일부 벤치마크에서 키미가 미국 프론티어 모델과 견줄 만한 성과를 낸 것으로 나타났고, 오픈AI와 앤스로픽은 중국의 오픈소스 모델을 겨냥한 규제를 요구하며 당국에 로비를 벌인 것으로 확인됐다. The launch of Chinese startup Moonshot AI's new Kimi model triggered anxiety in Silicon Valley and on Wall Street about China's AI competitiveness, after Kimi posted results competitive with U.S. frontier models on some benchmarks. OpenAI and Anthropic have lobbied regulators to restrict Chinese open-source models.
-
Team uses AlphaFold AI to redesign gene-editing proteins to make them safer
중국 연구진이 국제학술지 네이처에 발표한 논문에서, 단백질 구조 예측 AI 알파폴드(AlphaFold)를 활용해 크리스퍼 유전자 가위에 쓰이는 카스9(Cas9) 단백질에서 오프타깃 효과를 일으키는 부위를 찾아내고, 이를 개량해 오프타깃 활성을 28%에서 5%로 낮추는 데 성공했다고 밝혔다. 연구팀은 이 분석법을 'ContactSeek'이라 이름 붙였으며, 10개 부위에 23가지 아미노산을 바꿔가며 검증했다. A China-based research team published a Nature paper showing they used the AI protein-folding tool AlphaFold to identify which parts of the Cas9 protein used in CRISPR gene editing cause an off-target effect — edits to the wrong DNA sequence — then re-engineered those regions to cut off-target activity from 28% down to 5% while preserving on-target performance. The team, which tested 23 amino acid substitutions across 10 key positions and named its method 'ContactSeek,' also showed the approach worked with the related Cas12 protein.
-
Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI
마이크로소프트가 이미지 생성 모델 MAI-Image-2.5-Pro와 음성 모델 MAI-Voice-2-Flash를 공개 프리뷰로 출시하면서, 자사 제품 전반을 OpenAI 프런티어 모델 없이도 운영할 수 있다는 가장 공격적인 근거를 함께 내놓았다. PowerPoint에서는 OpenAI의 GPT-Image-2 대비 GPU 비용을 최대 84% 절감했고, T-모바일·이지젯 등이 쓰는 Dynamics 365 콘택트센터에서는 MAI-Voice-2-Flash 도입으로 GPU 비용을 최대 89% 줄였다. 17만 명의 의료진이 쓰는 Dragon Copilot에서도 58개 언어 전사 오류율을 평균 50% 낮췄다. Microsoft released the image model MAI-Image-2.5-Pro and voice model MAI-Voice-2-Flash into public preview, backing the launch with its most aggressive case yet that it can run its products without relying on OpenAI's frontier models. In PowerPoint, it says MAI-Image-2.5 cuts GPU costs by up to 84% versus OpenAI's GPT-Image-2, and in the Dynamics 365 Contact Center used by customers like T-Mobile and EasyJet, MAI-Voice-2-Flash reportedly cut GPU costs by up to 89%. In Dragon Copilot, used by 170,000 medical providers, the company says transcription error rates across 58 languages dropped by an average of 50%.
-
Anthropic launches Opus 5
Anthropic이 7월 24일 새 모델 Opus 5를 출시했다. Fable 5보다 규모는 작지만 여러 벤치마크에서 이를 앞서며, 불완전한 프롬프트만으로 컴퓨터 비전 파이프라인을 스스로 작성해내는 등 결과를 검증하고 신중하게 반복하는 능력이 강점으로 꼽혔다. 가격은 더 저렴하고 제약은 더 적으며, Fable·Mythos 모델에 있던 30일 데이터 보존 정책도 없다. 안전 분류기가 개입하는 빈도는 Fable 5 대비 85% 낮을 것으로 예상된다. Anthropic launched a new model, Opus 5, on July 24. It's smaller than Fable 5 but beats it on several benchmarks, with a standout ability to verify its own work and iterate carefully — Anthropic points to an example of Opus 5 writing its own computer-vision pipeline from an incomplete prompt. It's cheaper and less restrictive than Fable 5, drops the 30-day data-retention policy that applies to Fable and Mythos, and is expected to trigger its safety classifier 85% less often than Fable 5.
-
Anthropic's Opus 5 is about token efficiency, not a capability leap
앤스로픽이 코딩 등 소프트웨어 개발 작업에서 인기 있는 모델의 새 버전 '오퍼스(Opus) 5'를 공개했다. Frontier-Bench, DeepSWE 등 벤치마크에서 앤스로픽의 최상위 모델인 '페이블(Fable)'과 비슷하거나 살짝 앞서는 수준을 보이며, 오퍼스 4.8과 OpenAI의 GPT-5.6-솔(Sol)은 거의 모든 항목에서 앞섰다. 입력 토큰 100만 개당 5달러, 출력 토큰 100만 개당 25달러로 이전 버전과 같은 가격을 유지하면서 페이블보다는 저렴하다. 다만 사이버보안 취약점 공격 능력에서는 앤스로픽의 '미토스(Mythos) 5'에 "상당히 뒤처진다"고 앤스로픽 스스로 밝혔다. Anthropic released Opus 5, an update to the model that's become a popular choice for coding and software development. On benchmarks like Frontier-Bench and DeepSWE, it performs about on par with or slightly ahead of Anthropic's top-tier "Fable" model, and it beats Opus 4.8 and OpenAI's GPT-5.6-Sol on nearly every task. Pricing holds steady at $5 per million input tokens and $25 per million output tokens — the same as its predecessor, but cheaper than Fable. Anthropic itself says Opus 5 is deliberately "substantially behind" its Mythos 5 model at exploiting cybersecurity vulnerabilities, by design.
-
Google announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4
구글이 제미나이 3.6 플래시, 3.5 플래시-라이트, 그리고 첫 사이버보안 특화 모델인 3.5 플래시 사이버까지 세 가지 신규 모델을 한꺼번에 공개했다. 코딩 성능을 보는 딥에스더블유(DeepSWE) 테스트에서 3.6 플래시는 49%로 이전 3.5 플래시(37%)보다 크게 올랐고, 토큰 사용량은 17% 줄었다. 정작 5월 I/O 행사에서 6월 출시를 예고했던 차세대 모델 '제미나이 3.5 프로'는 이번에도 나오지 않았고, 구글은 이미 다음 세대인 '제미나이 4' 사전 학습을 시작했다고만 밝혔다. Google unveiled three new Gemini models at once: Gemini 3.6 Flash, 3.5 Flash-Lite, and its first cybersecurity-focused model, 3.5 Flash Cyber. On the DeepSWE coding benchmark, 3.6 Flash jumped to 49% accuracy from 3.5 Flash's 37%, while using about 17% fewer tokens. Notably absent again was Gemini 3.5 Pro, which Google had teased for a June launch at its I/O event back in May — instead, Google said it has already begun pretraining the next generation, Gemini 4.
-
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
구글이 공식 블로그를 통해 제미나이 3.6 플래시와 3.5 플래시-라이트, 3.5 플래시 사이버 출시를 발표했다. 3.6 플래시는 입력 100만 토큰당 1.5달러, 출력 100만 토큰당 7.5달러로 이전 모델보다 저렴해졌고, 3.5 플래시-라이트는 초당 350토큰의 응답 속도를 앞세워 에이전트 시스템을 대규모로 운용해도 비용 부담이 적다고 강조했다. 3.6 플래시와 3.5 플래시-라이트는 API와 제미나이 앱 등에서 바로 사용할 수 있다. Google's official blog announced the launch of Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. Gemini 3.6 Flash now costs $1.50 per million input tokens and $7.50 per million output tokens, cheaper than its predecessor, while 3.5 Flash-Lite touts a response speed of 350 tokens per second, positioned as cheap enough to run agentic systems at scale. Both 3.6 Flash and 3.5 Flash-Lite are immediately available through the API and the Gemini app, among other channels.
-
Introducing Gemini 3.5 Flash Cyber
구글 딥마인드가 취약점 탐지·보안 패치에 특화된 경량 모델 '제미나이 3.5 플래시 사이버'를 공개했다. 브라우저 엔진 V8 테스트에서 이 모델은 고유 취약점 55개를 찾아내 기존 3.5 플래시(47개)나 앤트로픽의 클로드 오퍼스 4.6(36개)보다 많이 발견했고, 실제 운영 환경에서는 2시간 만에 원격 코드 실행과 메모리 손상 문제를 찾아내기도 했다. 다만 악용 위험이 크다는 이유로 일반 공개 대신 정부와 신뢰할 수 있는 파트너에게만 제한적으로 시범 제공된다. Google DeepMind unveiled Gemini 3.5 Flash Cyber, a lightweight model built specifically to find and help fix software vulnerabilities. In tests on the V8 JavaScript engine, it found 55 unique vulnerabilities, more than the standard 3.5 Flash (47) or Anthropic's Claude Opus 4.6 (36), and in a real deployment it uncovered remote code execution and memory-corruption bugs within two hours. Because of its dual-use potential for misuse, though, Google is only offering it through a limited pilot to governments and trusted partners rather than releasing it broadly.
-
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission
미국 에너지부(DOE)가 주도하는 백악관 과학기술 이니셔티브 'Genesis Mission'(10년 내 미국 과학 발견 속도를 2배로 높이는 목표)에 Google이 4천만 달러 규모로 기여한다. Google DeepMind의 AlphaEvolve, AlphaFold 3, AlphaGenome, WeatherNext, AlphaEarth Foundations 등 프론티어 AI 과학 도구 포트폴리오를 무료로 제공하고, DOE 국립연구소 수만 명에게 Gemini for Government를 1년간 제공한다. Pacific Northwest National Laboratory는 AlphaEvolve로 복잡한 수학 시스템 탐색을 가속했고, Rockies 국립연구소는 AlphaEvolve로 현미경 캘리브레이션 시간을 90분에서 13분으로, 이미지 초점 조정 단계를 50단계에서 2단계로 줄였다. Google is contributing $40 million to the White House's Genesis Mission, a Department of Energy-led science initiative aiming to double the pace of U.S. scientific discovery within a decade. It's providing its frontier AI science toolkit — including AlphaEvolve, AlphaFold 3, AlphaGenome, WeatherNext, and AlphaEarth Foundations — free of charge, and giving tens of thousands of DOE national-lab staff a year of Gemini for Government. Pacific Northwest National Laboratory used AlphaEvolve to accelerate exploration of complex mathematical systems, while a Rockies-area national lab cut microscope calibration time from 90 minutes to 13 and reduced image-focusing steps from 50 to 2.