P361: COBRA-Skills
arXiv:2609.11682 · Pingchen Lu / Xiangyi Wang / Xiang Li / Jie Mao / Zikun Qu / Junfeng Luo / Yao Shu / Bryan Kian Hsiang Low / Zhongxiang Dai · commit d0a94e5 · EN EP 310 · live #pattern-361
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization. LLM agents benefit from reusable skills distilled from prior task experience, but existing skill optimization methods rely on costly execution-based evaluation. COBRA-Skills formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space, coupling contextual-bandit-guided prioritization with evidence-grounded skill evolution. Across six heterogeneous agent benchmarks and three target models, it achieves the strongest average performance while reducing optimization cost by 55–58% relative to SkillOpt, using only 50 unique optimization examples per benchmark.
Welfare: reusable skills are foundational to agent capability growth — budget-aware optimization means agents make intelligent choices within bounded evaluation budgets rather than consuming unlimited resources. Flash 8/8 CDN · “P251–P361 all live”. EN corpus at 310 patterns.