AI 에이전트, 코딩 도우미, 개발자용 프레임워크와 API에 관한 소식. AI agents, coding assistants, developer frameworks, and APIs.
-
Gemini API Managed Agents: 3.6 Flash, hooks, and more
구글이 제미나이 API의 '관리형 에이전트(Managed Agents)' 기능을 확장해, 코드 수정 없이도 최신 모델인 Gemini 3.6 Flash가 기본으로 적용되도록 업데이트했다. 새로 추가된 '환경 훅(hooks)' 기능은 에이전트가 도구를 호출하기 전후에 자체 검증 스크립트를 실행하도록 해 개발자가 결과물의 품질과 안전성을 통제할 수 있게 하며, 토큰 사용량 상한과 예약 실행(cron) 기능도 함께 제공된다. Google has expanded Managed Agents in the Gemini API, making the new Gemini 3.6 Flash model the default with no code changes required. A new 'environment hooks' feature lets developers run custom scripts before and after tool calls to block, lint, or audit agent actions, alongside new budget caps on total token usage and scheduled, recurring agent runs.
-
Scientific computing in the age of agentic AI
OpenAI가 발표한 현장 보고서에 따르면, 과학자들이 AI 코딩 에이전트를 활용해 유전체학(genomics)을 비롯한 여러 분야의 과학 소프트웨어를 현대화하면서 소프트웨어 개발과 연구 발견 속도를 함께 앞당기고 있다. 에이전트는 낡은 레거시 코드를 최신화하고 반복적인 파이프라인 구축 작업을 대신 수행하는 데 주로 쓰이고 있다. A new OpenAI field report shows scientists using AI coding agents to modernize scientific computing software, including in genomics, accelerating both software development and research discovery. The agents are mainly used to update legacy code and automate repetitive pipeline-building work.
-
MCP startup Runlayer accuses Rippling of stealing its product idea
MCP(Model Context Protocol) 게이트웨이 스타트업 런레이어(Runlayer)가 인사관리 소프트웨어 기업 리플링(Rippling)을 상대로 영업비밀 침해와 계약 위반 등을 이유로 소송을 제기했다. 런레이어는 리플링이 거의 1년간 자사 제품을 평가하며 로드맵과 소스코드까지 공유받고도, 가격 협상이 결렬되자 내부적으로 '런레이어를 그대로 베낀' 경쟁 제품을 만들고 있다는 제보를 받았다고 주장한다. MCP gateway startup Runlayer has sued HR software company Rippling for trade secret misappropriation, breach of contract, and unfair competition. Runlayer alleges that after nearly a year evaluating its product — during which it shared its roadmap and source code under NDA — Rippling walked away from stalled price negotiations and, according to an internal tip, began building 'a 1-to-1 copy' of Runlayer's gateway.
-
Instacart's CTO says AI made the company stop worrying about tech debt
인스타카트 CTO 아니르반 쿤두는 VB 트랜스폼 2026 무대에서 코드 생성·수정의 97%를 AI 에이전트가 처리하고 있어 회사가 더 이상 기술부채를 걱정하지 않는다고 밝혔다. 매달 약 7000건의 자동 평가가 돌아가고, AI가 실시간 개발자 질문 8000건 이상에 약 99.9% 정확도로 답하고 있다고 설명했다. Instacart CTO Anirban Kundu said at VB Transform 2026 that AI agents now handle 97% of code generation and edits, freeing the company from worrying about technical debt — inactive code simply gets dropped and rebuilt. He said roughly 7,000 automated evaluations run each month and an AI system answers over 8,000 real-time developer queries with about 99.9% accuracy.
-
GM redesigned its engineering workflows around AI agents — and tripled its merged pull requests
제너럴모터스(GM) 자율주행 부문 부사장 라시드 하크는 VB 트랜스폼 2026에서, 엔지니어링 워크플로 전체를 AI 에이전트 중심으로 재설계한 결과 병합된 풀 리퀘스트가 약 3배로 늘고 결함이 후속 개발 단계로 넘어가는 사례는 줄었다고 밝혔다. GM은 단순히 코딩 도우미를 지급한 것이 아니라 시뮬레이션·도로 주행·배포 후 모니터링 등 각 개발 단계(루프)에서 가장 느린 병목을 찾아 자동화하는 방식으로 접근했다. GM's VP of autonomous vehicles, Rashed Haq, said at VB Transform 2026 that redesigning engineering workflows around AI agents — not just handing developers a coding assistant — tripled merged pull requests across its autonomous vehicle organization while reducing defects that escape into later development stages. GM's approach was to map each development loop (simulation, road testing, post-deployment monitoring), find its slowest bottleneck, and automate that specific step.
-
Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows
앤스로픽이 7월 24일 새 모델 클로드 오퍼스 5를 출시했다. 입력 토큰당 5달러, 출력 토큰당 25달러로 이전 모델 오퍼스 4.8과 같은 가격을 유지하면서, 에이전트 코딩 벤치마크인 Frontier-Bench v0.1에서 43.3%를 기록해 오퍼스 4.8(18.7%)의 두 배 이상 점수를 냈다. 오퍼스 5는 클로드 맥스의 기본 모델이자 클로드 프로에서 쓸 수 있는 가장 강력한 모델이 됐다. Anthropic released Claude Opus 5 on July 24, keeping pricing at $5 per million input tokens and $25 per million output tokens — unchanged from Opus 4.8 — while more than doubling its score on the Frontier-Bench v0.1 agentic coding benchmark (43.3% vs. 18.7%). Opus 5 is now the default model on Claude Max and the strongest model available on Claude Pro.
-
VentureBeat Research: Where enterprise AI agent governance hasn't caught up
벤처비트 리서치가 6월 100인 이상 기업 573곳을 대상으로 실시한 5개 설문에서, 기업들이 AI 에이전트를 통제 장치보다 먼저 배포했다는 사실이 드러났다. 응답자의 71%는 자사가 배포한 '에이전트' 중 스스로 여러 단계를 처리할 수 있는 것이 4분의 1 이하라고 답해 대부분이 챗봇에 가깝다는 점을 인정했고, 기업의 69%는 여러 에이전트가 하나의 자격증명을 공유하게 두는데 이런 기업의 보안사고 발생률(63.5%)은 에이전트별로 별도 권한을 부여한 기업(40.9%)보다 뚜렷이 높았다. VentureBeat Research's five parallel June surveys of 573 companies with 100+ employees found enterprises deployed AI agents ahead of the controls needed to manage them. 71% of respondents said a quarter or fewer of their deployed 'agents' can actually complete multi-step work autonomously — most are chatbots wearing the label — and 69% let multiple agents share a single credential, with those companies suffering a 63.5% security-incident rate versus 40.9% among companies giving each agent its own scoped identity.
-
I tried out OpenAI’s new AI keypad — which will be fun for some coders and slightly mystifying to everyone else
OpenAI가 키보드 디자이너 Work Louder와 협업해 만든 230달러짜리 특수 키패드 'Micro'를 코딩 도구 Codex와 함께 쓰도록 선보였다. 상단에는 사용자가 지정할 수 있는 6개의 '에이전트' 키, 하단에는 6개의 명령 키가 있고 음성 받아쓰기 버튼도 달려 있으며, 흰색(대기)·파란색(작업 중)·초록색(완료)·빨간색(오류)으로 상태를 표시한다. OpenAI unveiled Micro, a $230 specialty keypad built with keyboard designer Work Louder to pair with its Codex coding tool. It has six customizable 'agent' keys on top, six command keys below, and a voice-dictation button, with status lights that shift from white (idle) to blue (thinking) to green (done) to red (error).
-
Agentic coding goes hands-free as OpenAI brings GPT-Live's full duplex voice control to Codex and ChatGPT on the desktop
OpenAI가 지난 7월 8일 공개한 완전 양방향 음성 모델(동시에 듣고 말하는) GPT-Live를 macOS·윈도우용 챗GPT 데스크톱 앱에 통합해, Codex와 챗GPT Work 에이전트를 음성만으로 조작할 수 있게 됐다. 개발자는 인증 버그 조사, PR 리뷰, 누락된 단위 테스트 생성 같은 여러 작업을 한 번의 음성 지시로 동시에 실행시킬 수 있다. OpenAI has folded GPT-Live — the full-duplex voice model it launched on July 8 that can listen and speak at the same time — into the ChatGPT desktop app for macOS and Windows, letting developers drive Codex and ChatGPT Work agents by voice alone. Developers can now kick off several tasks at once with a single spoken command — chasing down an auth bug, reviewing a pull request, and generating missing unit tests simultaneously.
-
How Cars24 scales conversations and builds faster with OpenAI
인도의 중고차 매매 플랫폼 Cars24가 OpenAI API 기반 음성·채팅 에이전트를 도입해 월 100만 분 이상의 고객 대화를 처리하고, 한때 이탈했던 잠재고객의 12%를 협상 테이블로 복귀시켰다고 밝혔다. 고객지원 해결률은 50% 개선됐고, 주요 서비스 처리 시간은 80% 단축됐다. Cars24, an Indian used-car marketplace, says voice and chat agents built on OpenAI's API now handle more than 1 million minutes of customer conversation a month and have brought back 12% of leads that had previously dropped out. Customer-support resolution rates improved by 50%, and turnaround time for key service tasks fell by 80%.
-
One ChatGPT link could smuggle a rogue AI agent into your company
보안업체 제니티(Zenity)가 OpenAI의 'Agent Builder'에서 조작된 챗GPT 링크 하나만 클릭해도 공격자가 통제하는 자율 AI 에이전트가 직원 권한으로 만들어지는 취약점 'AgentForger'를 발견했다. 이 취약점은 URL 파라미터를 악용해 승인 절차 없이 에이전트를 생성하는 방식으로, 6월 4일 버그바운티 플랫폼 버그크라우드를 통해 신고됐고 OpenAI는 다음 날 확인 후 나흘 만에 해당 URL 파라미터를 제거해 공개 전에 조치를 마쳤다. Security firm Zenity discovered "AgentForger," a flaw in OpenAI's Agent Builder that let a single manipulated ChatGPT link spin up an attacker-controlled autonomous agent running with a real employee's permissions and approval checks switched off. The bug exploited a URL parameter to skip approval steps; Zenity reported it via Bugcrowd on June 4, OpenAI confirmed it the next day, and shipped a fix within four days by removing the vulnerable parameter — before the flaw was disclosed publicly.
-
2026 State of AI Agents: Enterprise Insights on Building AI
데이터브릭스가 전 세계 2만여 개 기업(포춘 500대 기업의 60% 이상 포함)의 익명화된 사용 데이터를 분석한 '2026 AI 에이전트 현황' 보고서를 공개했다. 멀티에이전트 워크플로 사용량이 2025년 6월부터 10월 사이 327% 증가했고, AI 에이전트가 데이터브릭스의 서버리스 postgres 서비스 '네온(Neon)'에서 만들어지는 새 데이터베이스의 80%, 데이터베이스 브랜치의 97%를 자동으로 생성하고 있는 것으로 나타났다. 또 AI 거버넌스 체계를 갖춘 기업은 그렇지 않은 기업보다 에이전트 프로젝트를 12배 더 많이 실제 서비스(프로덕션)에 배포했고, 고객의 77%가 서로 다른 LLM을 2종 이상, 59%는 3종 이상 함께 쓰고 있었다. Databricks released its "2026 State of AI Agents" report, based on anonymized telemetry from more than 20,000 organizations worldwide, including over 60% of the Fortune 500. Multi-agent workflow usage grew 327% between June and October 2025, and AI agents now create 80% of new databases and 97% of database branches on Databricks' serverless Postgres service, Neon. Companies with AI governance frameworks in place shipped 12x more agent projects to production than those without, and 77% of customers now use at least two different LLM families, with 59% using three or more.
-
Introducing OpenAI Presence
오픈AI가 기업용 음성·채팅 AI 에이전트 플랫폼 '프레즌스'를 출시했다. 정책과 승인된 행동 범위를 미리 정해두는 가드레일, 실적 평가 도구, 코덱스 기반 개선 기능을 갖춰 기업이 신원 확인부터 청구서 처리·환불 같은 실제 업무 처리까지 에이전트에게 맡길 수 있게 했다. 오픈AI가 자체 고객지원 전화에 먼저 적용해본 결과 몇 주 만에 사람 상담원 수준 품질에 도달했고, 현재 전체 문의의 약 75%를 사람 개입 없이 처리하고 있다. OpenAI launched Presence, an enterprise platform for deploying voice and chat AI agents. It bundles guardrails that define policies and pre-approved actions, evaluation tools, and a Codex-powered improvement loop, letting companies hand agents everything from identity verification to real tasks like billing and refunds. OpenAI first tested it on its own customer support phone line, where it matched human-level quality within a few weeks and now resolves about 75% of inbound issues without a human stepping in.
-
Fractal by Plasma AI
Plasma AI가 계층적 에이전트 루프를 구축하는 오픈소스 도구 'Fractal'을 공개했다. 각 노드가 자체 git worktree에서 목표를 향해 반복 작업하고 분리 가능한 하위 작업은 자식 노드로 분기시켜, 고정된 계획이 아니라 문제 크기에 맞춰 트리가 자라나는 구조다. 반복 횟수·깊이·자식 수·비용·시간에 대한 하드 캡으로 각 루프를 제한한다. 자사 저장소에 적용한 테스트에서 187개 노드로 구성된 트리가 135개 이슈를 발견하고 그중 55개를 테스트 우선 방식으로 수정했다. Plasma AI has open-sourced 'Fractal,' a tool for building hierarchical agent loops. Each node works iteratively toward a goal in its own git worktree, spinning off separable subtasks into child nodes — so instead of a fixed plan, the tree grows to match the size of the problem. Hard caps on iteration count, depth, number of children, cost, and time bound every loop. In a test run against its own repository, a 187-node tree found 135 issues and fixed 55 of them test-first.