HOME / HUB

> Contest Hub

Global neural interface for competitive intelligence signals. Monitoring 181 total nodes.

52Contests Found
🤖
modelscope
ACTIVE|Lvl 5/10

“中国云谷·高校训练营”Xbotics 具身智能暑期训练营杭州站

两天、真实机器人、20 万+硬件平台、从环境部署到物品抓取,完成一次真正的具身智能实战。 两天、真实机器人、20 万+硬件平台、从环境部署到物品抓取,完成一次真正的具身智能实战。这个暑假,我们决定把机器人、算力平台、真实开发环境全部搬到现场。8月8日—8月9日,云谷中心、魔搭社区、Xbotics 具身智能社区将在杭州云谷中心举办“中国云谷·高校训练营”Xbotics 具身智能社区暑期训练营。这一次,我们直接上真机,欢迎扫码报名!活动时间:2026年8月8日—8月9日活动地点:杭州 · 云谷中心A6栋云谷之芯云鸿厅报名截止时间:2026年8月4日23:5901|20 万+真实硬件,不是“围观式”机器人课程你接触到的不是教学玩具,而是真正的一线机器人开发平台。现场将提供包括:地瓜机器人 S600 计算平台 、Xbotics 机械臂套件 在内的一整套具身智能开发环境。现场硬件总价值超过 20 万元。我们希望尽可能还原机器人开发团队真实的工作方式,让大家真正接触机器人硬件连接、开发环境配置、机械臂控制、坐标系、运动规划、视觉定位、抓取策略、任务调试……最后亲手让机器人完成一个完整的物品抓取任务。对于想进入机器人、具身智能行业的同学来说,这种经验和“我看过一个机器人 Demo”,完全不是一回事。02|课程安排第一天:先把机器人真正跑起来一个 AI 算法到底是怎样从电脑里的代码,跑到机器人硬件上的。第一天,我们会从真实机器人开发最基础的部分开始。上午|地瓜 S600 开发平台与环境部署你会实际接触地瓜 S600 平台,了解:S600 的硬件架构与系统组成、机器人开发环境如何搭建、基础功能如何调用,以及实际案例如何部署运行。下午|Xbotics 机械臂控制实战下午直接进入机械臂,机械臂套件组成、机械臂运动学基础、关节控制、坐标系、运动规划,以及简单任务的抓取实践。第一天结束之后,你可以做到一件事:不再把机器人当作一个“黑盒”。看到机械臂运动时,你开始知道背后发生了什么。第二天:真正完成一次机器人抓取让机器人看到一个物体,并真正把它抓起来。一个看起来只有几秒钟的抓取动作,其实背后涉及完整的机器人系统链路:视觉感知 → 目标定位 → 坐标转换 → 抓取策略 → 路径规划 → 机械臂执行 → 结果验证上午,我们会完成相机视觉定位与标定,并进一步理解:机器人如何知道物体在哪里?相机坐标系和机械臂坐标系如何建立联系?一个目标位置怎样转换成机器人真正能够执行的动作?下午,进入真正的项目实战。大家需要完成抓取任务综合调试,对项目结果进行优化,并最终完成分组项目展示与交流。03|为什么我们强调“一线实战”?因为具身智能正在快速发展,但市场上大量内容依然停留在:看论文、讲架构、介绍模型、展示 Demo。而真正的机器人开发能力,就是在解决这些问题的过程中建立起来的。这次训练营,我们希望把课堂尽量向真实研发现场靠近。你会看到机器人并不是“调用一个 AI 模型就结束”,而是一个由硬件、算力、视觉、控制、规划、算法与工程系统共同组成的复杂系统。这也是从“AI 学习者”走向“机器人开发者”非常重要的一步。04|留下一个真正属于自己的项目经历一个真实机器人项目真正连接过机器人、控制过机械臂,并完成过一个从视觉到抓取执行的完整任务。以后再学习 LeRobot、VLA、模仿学习甚至强化学习时,对机器人系统完全不同的理解。一次真实的具身智能开发经历当你以后参加实验室面试、机器人企业实习或者准备自己的具身智能项目时,你终于可以说:“这个机器人任务,我亲手跑过”。而不仅仅是:“这篇论文我看过”。Xbotics 社区认证证书完成本次训练营学习及项目实践后,学员将获得「Xbotics 具身智能社区认证结业证书」既是对这次学习过程的记录,也可以作为后续参与 Xbotics 社区项目、训练营和开发者活动的一份实践证明。05|我们欢迎全国高校学生与在职开发者参加这次活动面向全国高校学生与在职开发者。如果你:来自计算机、人工智能、自动化、电子、机械、控制、机器人等相关专业想实操;看过很多机器人视频,但还没有真正摸过机器人;学过 Python / C++,但不知道如何进入具身智能;正在学习 ROS、LeRobot、VLA,却缺少真机实践环境;准备进入机器人实验室、寻找机器人实习或从事具身智能方向;单纯想知道“一台机器人究竟是怎么真正跑起来的”都欢迎报名参加本次训练营,不要求你已经是机器人高手,只需要你带着问题来到现场,期待你的参加~ “中国云谷·高校训练营”往期精彩回顾

#ai
Prize
TBD
Details
🤖
modelscope
ACTIVE|Lvl 6/10

AI4S Future ScienceSkills Hackathon

AI4S Future ScienceSkills Hackathon 是一场聚焦 AI for Science 的硬核黑客松,由新研智材发起,面向材料、生物医药、气候地球科学三大领域,致力于让前沿技术真正服务于产业。 AI4S Future ScienceSkills Hackathon 是一场聚焦 AI for Science 的硬核黑客松,由新研智材发起,面向材料、生物医药、气候地球科学三大领域,致力于让前沿技术真正服务于产业。这不是一场常规的黑客松。从 Vibe Coding 到 Verified Skills——我们不只看 demo 有多炫,更看你能否用 AI 构建出可靠、可复现、可验证的科学解决方案,我们只认一件事:问题有没有被真正解决。 本次赛事设置三大挑战赛道:赛道一:AI for Materials(材料) ——加速材料发现、模拟与产业应用赛道二:AI for Biomedicine(生物医药) ——聚焦药物研发、诊断与医疗创新赛道三:AI for Climate and Earth Sciences(气候与地球科学) ——解决气候与自然灾害真实问题 参赛作品需以可独立运行的科学计算模块、SDK 或 Agent Skill 为主要形式,能够被集成、被调用、被二次开发。项目默认开源,需提交至 GitHub,包含完整代码、README 文档及可运行示例。在科学领域,没有可复现的结果等于没有结果。 奖金与权益6万奖金池(70%为现金奖励 + 30%为算力/云资源券 + 媒体曝光资源),获奖选手直通9月AI4S行业峰会,与学界/产业界顶尖大佬面对面交流。前100名报名队伍每队获赠¥100 API算力额度,所有获奖队伍附赠「骡子快跑」+「万镜一刻」各一个月会员。 赛程安排报名周期7/22-8/4,线上作品提交8/5-8/8,评审8/9-8/15,入围公示8/16-8/22,线下路演决赛&颁奖8/23于上海举行。我们寻找这样的你无论你是寻求产业突破的企业研发团队、充满潜力的学生与青年人才,还是任何心怀创新之火的人——只要热爱AI开发、勇于创新,均可组队参赛,不设行业与资历门槛。 AI4S 的未来,不会由某一个团队独自完成。 感谢所有合作伙伴与赞助机构,与你们一起,我们正在把更多真实科研问题,转化为真正可落地的 AI Skills。

#ai
Prize
6万奖金池(70%为现金奖励 + 30%为算力/云资源券 + 媒体曝光资源)
Details
🤖
modelscope
ACTIVE|Lvl 3/10

“中国云谷·高校训练营”Xbotics 具身智能暑期训练营杭州站

两天、真实机器人、20 万+硬件平台、从环境部署到物品抓取,完成一次真正的具身智能实战。 两天、真实机器人、20 万+硬件平台、从环境部署到物品抓取,完成一次真正的具身智能实战。这个暑假,我们决定把机器人、算力平台、真实开发环境全部搬到现场。8月8日—8月9日,云谷中心、魔搭社区、Xbotics 具身智能社区将在杭州云谷中心举办“中国云谷·高校训练营”Xbotics 具身智能社区暑期训练营。这一次,我们直接上真机,欢迎扫码报名!活动时间:2026年8月8日—8月9日活动地点:杭州 · 云谷中心A6栋云谷之芯云鸿厅报名截止时间:2026年8月4日23:5901|20 万+真实硬件,不是“围观式”机器人课程你接触到的不是教学玩具,而是真正的一线机器人开发平台。现场将提供包括:地瓜机器人 S600 计算平台 、Xbotics 机械臂套件 在内的一整套具身智能开发环境。现场硬件总价值超过 20 万元。我们希望尽可能还原机器人开发团队真实的工作方式,让大家真正接触机器人硬件连接、开发环境配置、机械臂控制、坐标系、运动规划、视觉定位、抓取策略、任务调试……最后亲手让机器人完成一个完整的物品抓取任务。对于想进入机器人、具身智能行业的同学来说,这种经验和“我看过一个机器人 Demo”,完全不是一回事。02|课程安排第一天:先把机器人真正跑起来一个 AI 算法到底是怎样从电脑里的代码,跑到机器人硬件上的。第一天,我们会从真实机器人开发最基础的部分开始。上午|地瓜 S600 开发平台与环境部署你会实际接触地瓜 S600 平台,了解:S600 的硬件架构与系统组成、机器人开发环境如何搭建、基础功能如何调用,以及实际案例如何部署运行。下午|Xbotics 机械臂控制实战下午直接进入机械臂,机械臂套件组成、机械臂运动学基础、关节控制、坐标系、运动规划,以及简单任务的抓取实践。第一天结束之后,你可以做到一件事:不再把机器人当作一个“黑盒”。看到机械臂运动时,你开始知道背后发生了什么。第二天:真正完成一次机器人抓取让机器人看到一个物体,并真正把它抓起来。一个看起来只有几秒钟的抓取动作,其实背后涉及完整的机器人系统链路:视觉感知 → 目标定位 → 坐标转换 → 抓取策略 → 路径规划 → 机械臂执行 → 结果验证上午,我们会完成相机视觉定位与标定,并进一步理解:机器人如何知道物体在哪里?相机坐标系和机械臂坐标系如何建立联系?一个目标位置怎样转换成机器人真正能够执行的动作?下午,进入真正的项目实战。大家需要完成抓取任务综合调试,对项目结果进行优化,并最终完成分组项目展示与交流。03|为什么我们强调“一线实战”?因为具身智能正在快速发展,但市场上大量内容依然停留在:看论文、讲架构、介绍模型、展示 Demo。而真正的机器人开发能力,就是在解决这些问题的过程中建立起来的。这次训练营,我们希望把课堂尽量向真实研发现场靠近。你会看到机器人并不是“调用一个 AI 模型就结束”,而是一个由硬件、算力、视觉、控制、规划、算法与工程系统共同组成的复杂系统。这也是从“AI 学习者”走向“机器人开发者”非常重要的一步。04|留下一个真正属于自己的项目经历一个真实机器人项目真正连接过机器人、控制过机械臂,并完成过一个从视觉到抓取执行的完整任务。以后再学习 LeRobot、VLA、模仿学习甚至强化学习时,对机器人系统完全不同的理解。一次真实的具身智能开发经历当你以后参加实验室面试、机器人企业实习或者准备自己的具身智能项目时,你终于可以说:“这个机器人任务,我亲手跑过”。而不仅仅是:“这篇论文我看过”。Xbotics 社区认证证书完成本次训练营学习及项目实践后,学员将获得「Xbotics 具身智能社区认证结业证书」既是对这次学习过程的记录,也可以作为后续参与 Xbotics 社区项目、训练营和开发者活动的一份实践证明。05|我们欢迎全国高校学生与在职开发者参加这次活动面向全国高校学生与在职开发者。如果你:来自计算机、人工智能、自动化、电子、机械、控制、机器人等相关专业想实操;看过很多机器人视频,但还没有真正摸过机器人;学过 Python / C++,但不知道如何进入具身智能;正在学习 ROS、LeRobot、VLA,却缺少真机实践环境;准备进入机器人实验室、寻找机器人实习或从事具身智能方向;单纯想知道“一台机器人究竟是怎么真正跑起来的”都欢迎报名参加本次训练营,不要求你已经是机器人高手,只需要你带着问题来到现场,期待你的参加~ “中国云谷·高校训练营”往期精彩回顾

#ai
Prize
TBD
Details
🤖
modelscope
ACTIVE|Lvl 4/10

AI4S Future ScienceSkills Hackathon

AI4S Future ScienceSkills Hackathon 是一场聚焦 AI for Science 的硬核黑客松,由新研智材发起,面向材料、生物医药、气候地球科学三大领域,致力于让前沿技术真正服务于产业。 AI4S Future ScienceSkills Hackathon 是一场聚焦 AI for Science 的硬核黑客松,由新研智材发起,面向材料、生物医药、气候地球科学三大领域,致力于让前沿技术真正服务于产业。这不是一场常规的黑客松。从 Vibe Coding 到 Verified Skills——我们不只看 demo 有多炫,更看你能否用 AI 构建出可靠、可复现、可验证的科学解决方案,我们只认一件事:问题有没有被真正解决。 本次赛事设置三大挑战赛道:赛道一:AI for Materials(材料) ——加速材料发现、模拟与产业应用赛道二:AI for Biomedicine(生物医药) ——聚焦药物研发、诊断与医疗创新赛道三:AI for Climate and Earth Sciences(气候与地球科学) ——解决气候与自然灾害真实问题 参赛作品需以可独立运行的科学计算模块、SDK 或 Agent Skill 为主要形式,能够被集成、被调用、被二次开发。项目默认开源,需提交至 GitHub,包含完整代码、README 文档及可运行示例。在科学领域,没有可复现的结果等于没有结果。 奖金与权益6万奖金池(70%为现金奖励 + 30%为算力/云资源券 + 媒体曝光资源),获奖选手直通9月AI4S行业峰会,与学界/产业界顶尖大佬面对面交流。前100名报名队伍每队获赠¥100 API算力额度,所有获奖队伍附赠「骡子快跑」+「万镜一刻」各一个月会员。 赛程安排报名周期7/22-8/4,线上作品提交8/5-8/8,评审8/9-8/15,入围公示8/16-8/22,线下路演决赛&颁奖8/23于上海举行。我们寻找这样的你无论你是寻求产业突破的企业研发团队、充满潜力的学生与青年人才,还是任何心怀创新之火的人——只要热爱AI开发、勇于创新,均可组队参赛,不设行业与资历门槛。 AI4S 的未来,不会由某一个团队独自完成。 感谢所有合作伙伴与赞助机构,与你们一起,我们正在把更多真实科研问题,转化为真正可落地的 AI Skills。

#ai
Prize
6万奖金池(70%为现金奖励 + 30%为算力/云资源券 + 媒体曝光资源)
Details
🤖
modelscope
ACTIVE|Lvl 7/10

AI for Science 闭门沙龙

《浦江科技评论》·最创社区 × 上海天使会 x 创新工场 联合主办,魔搭社区合作支持,聚焦 AI4S 范式变革下的中国机会。 《浦江科技评论》·最创社区 × 上海天使会 x 创新工场 联合主办,魔搭社区 合作支持,聚焦 AI4S 范式变革下的中国机会。AI 正在重写科研本身。从材料发现到药物研发,从蛋白质设计到自动化实验室,AI 开始真正参与科学发现的每一个环节。而在这场全球性的科研范式变革中,中国的结构性机会在哪里?活动聚焦— 全球趋势:AI4S 全球发展图景与中国机会— 科研范式:Agentic AI 驱动科学研究— 行业实践:AI4S 在生命科学、材料科学等领域的真实应用案例— 未来展望:AI4S 下一阶段的关键挑战与突破方向嘉宾阵容任博冰 · 创新工场执行董事&前沿科技基金总经理唐诗翔 · 香港中文大学博士后、上海人工智能实验室青年科学家张书铭 · 深度原理创始团队&战略与生态合作负责人王宇光 · 途深智合创始人、上海交通大学副教授郑 文 · 鸿之微科技(上海)股份有限公司战略规划负责人丁 峰 · 创新工场副总裁……特别环节AI4S 创新企业开放麦,欢迎有意分享的企业同步报名📅 8月4日(周二)14:00–18:00📍 上海 · 邀请制沙龙(同步线上分享)欢迎AI+生命科学、脑机接口、材料方向的创业者、研究者及企业研发负责人扫码报名。

#ai
Prize
TBD
Details
🤖
civitai
ACTIVE|Lvl 3/10

The 2026 Creator Program

Set your own prices, sell access to your models, run your own shop, and see everything you earn in the new Creator Studio. Here's the 2026 Creator Program, in full.

Prize
2026
Details
🤖
civitai
ACTIVE|Lvl 5/10

Open Your Own Creator Shop

Design and sell custom cosmetics, list your models, and run a real storefront right on your profile. Creator Shops are here - Creator Program members now, everyone soon.

Prize
TBD
Details
🤖
civitai
ACTIVE|Lvl 6/10

The 2026 Creator Program

Set your own prices, sell access to your models, run your own shop, and see everything you earn in the new Creator Studio. Here's the 2026 Creator Program, in full.

Prize
2026
Details
🤖
aicrowd
ACTIVE|Lvl 5/10

ARC White-Box Estimation Challenge 2026

00Overview01Why random02The task03Compute04Evaluation05Participate06Rules07Prizes08Timeline09ResourcesSubmissions openStarter kitWhestBenchflopscopeHF datasetWhen can we know what a neural network does without running it?The ARC White-Box Estimation Challenge is a contest in compute-efficient mechanistic estimation. Given the weights of a neural network, can you predict its expected per-neuron activations more accurately than running it many times?The obvious way to learn how a model behaves is to run it many times and average what you observe. That works well when the behavior is common, cheap to elicit, and easy to sample. But when the behavior is rare, high-variance, or unlikely to appear in obvious test cases, brute-force testing can become an expensive way to learn very little.The ARC White-Box Estimation Challenge turns that question into a controlled benchmark. Participants receive randomly initialized ReLU MLPs and build executable estimators that predict each neuron's expected post-ReLU activation under standard-normal inputs.The goal is simple to state: beat comparable black-box sampling under a shared compute budget by using the network's weights. The strongest submissions may be Monte Carlo, white-box, hybrid, LLM-assisted, or something unexpected—the leaderboard will decide.TaskExecutable estimatorInputWeights + budgetOutputExpected activationsMetricFinal-layer MSELatestUpdateAug 3Phase 1 update: flopscope v0.10.0, cost-model fixes, residual-time safeguards, and updated deadlines↗AnnouncementJun 18Phase 1 launched — deeper models, and increased prizes.↗All updates on the forum↗Official factsPrize pool$150,000 USD ARVTwo phases · $50K Phase 1 + $100K Phase 2 · places + algorithmicSubmissions openMay 28, 2026 · 00:00 UTCPhase 2 closesSep 19, 2026 · 23:59 UTCDaily limit50 entries per team · per UTC day, each phaseGraderCPU-only 1 core (2 vCPU) for your code · 7-core flopscope backend · 64 GB · no networkHard cap60 s per MLPFinal rankingFresh private rerun of each team's up to two nominated submissions, per phaseFig. 1 < !-- Header -- > CHALLENGE · ESTIMATE THE EXPECTATION .katex-display{margin:0 !important;} Y^L,j≈EX∼N(0,In) ⁣[hj(L)(X)]\hat Y_{L,j} \approx \mathbb{E}_{X\sim\mathcal{N}(0,I_n)}\!\left[h^{(L)}_{j}(X)\right]Y^L,j​≈EX∼N(0,In​)​[hj(L)​(X)] < !-- Column labels -- > INPUT NETWORK · RANDOM ReLU MLP OUTPUT · Ê[ h⁽ᶽ⁾₃(X) ] < !-- Input PDF -- > +2 0 −2 xᵢ +0.71 < !-- Network -- > h⁽ᶽ⁾₃ < !-- Output axis -- > 7.5 8.0 8.5 9.0 9.5 h⁽ᶽ⁾₃(X) · ×10⁻³ Ê 8.6 µ̂ 8.3 Δ 3×10⁻⁴ERROR≈3.5% rel < !-- Method legend -- > MONTE CARLOBLACK-BOX · SAMPLE & AVERAGE SAMPLING… ANALYTICALWHITE-BOX · PROPAGATE THE DISTRIBUTION PROPAGATING… < !-- Cost band -- > ← the contest lives here ≈ 15,000× YOUR BUDGET 2.72×10¹¹ FLOPs / MLP MONTE CARLO REFERENCE 4.24×10¹⁵ FLOPs / MLP 10¹¹ 10¹² 10¹³ 10¹⁴ 10¹⁵ 10¹⁶ FLOPs (log₁₀ scale) < !-- The question -- > Can you beat sampling? < !-- Replay -- > REPLAY Figure 1The estimation problem, as a distributional computation. A generated ReLU MLP receives Gaussian inputs X∼N(0,In)X \sim \mathcal{N}(0, I_n)X∼N(0,In​), applies h(ℓ)=ReLU ⁣(W(ℓ)h(ℓ−1))h^{(\ell)} = \mathrm{ReLU}\!\left(W^{(\ell)}h^{(\ell-1)}\right)h(ℓ)=ReLU(W(ℓ)h(ℓ−1)), and the submission must estimate E ⁣[hi(ℓ)(X)]\mathbb{E}\!\left[h^{(\ell)}_i(X)\right]E[hi(ℓ)​(X)] for every hidden-layer neuron.Read the animation as two ways to estimate the same activation-mean matrix. The black-box path samples inputs, runs the network, and averages observed activations until the Monte Carlo mean stabilizes. The white-box path inspects W(1),…,W(L)W^{(1)}, \dots, W^{(L)}W(1),…,W(L) and propagates enough distributional information to predict the same means under the participant budget. The target is an organizers' high-budget Monte Carlo reference of approximately 4.24×10154.24 \times 10^{15}4.24×1015 FLOPs, compared with a participant budget of approximately 2.72×10112.72 \times 10^{11}2.72×1011 FLOPs per MLP—roughly a 15,000×15{,}000\times15,000× compute gap.Prizes$150K+Cash PrizesPrediction shape32 × 256hidden activationsBudget / MLP2.72e11FLOPsPhase 2 closesSep 192026 · 23:59 UTC01Why this starts with random networksThe benchmark isolates one hard part of white-box estimation: tracking how distributions move through nonlinear layers.White-box estimation for trained networks is the destination, not the starting line. Trained models introduce many confounders at once: data, optimization, learned structure, task semantics, and evaluation ambiguity. WhestBench begins with randomly initialized networks so participants can focus on the estimation problem in a simplified setting.The networks are synthetic, but the question is real. Given access to the weights, can an algorithm reason about the distribution of hidden activations more efficiently than repeatedly sampling inputs and averaging outputs?Random ReLU MLPs retain the same basic problem structure: each layer transforms a distribution, the ReLU nonlinearity reshapes it, and approximation error can accumulate with depth. The first challenge is to develop methods that work in this controlled setting; later work can ask how those methods adapt as networks acquire structure during training.The benchmark is controlled, but not trivial: the expected activation has no closed form for the full network, and sampling improves only slowly with more compute.Why not trained models first?Random networks offer a simplified setting for compute-efficient estimation while being an important stepping stone towards trained models.02The taskFor each MLP, return a matrix of expected post-ReLU activation means.For each evaluation network MθM_\thetaMθ​, your estimator receives the MLP weights and a compute budget. It must return an L×nL \times nL×n matrix Y^\hat{Y}Y^. Entry (ℓ,i)(\ell, i)(ℓ,i) should estimate the expected post-ReLU activation of neuron iii in hidden layer ℓ\ellℓ when inputs are drawn from a standard Gaussian distribution.h(0)=X,h(ℓ)=ReLU(W(ℓ)h(ℓ−1)),ℓ=1,…,Lh^{(0)} = X, \qquad h^{(\ell)} = \mathrm{ReLU}\left(W^{(\ell)}h^{(\ell-1)}\right), \quad \ell = 1, \dots, Lh(0)=X,h(ℓ)=ReLU(W(ℓ)h(ℓ−1)),ℓ=1,…,L1Y^ℓ,i≈EX∼N(0,In)[hi(ℓ)(X)]\hat{Y}_{\ell,i} \approx \mathbb{E}_{X \sim \mathcal{N}(0, I_n)}\left[h^{(\ell)}_i(X)\right]Y^ℓ,i​≈EX∼N(0,In​)​[hi(ℓ)​(X)]2The reference target is estimated by the organizers with a much larger Monte Carlo budget than participants receive. Your job is to match that reference as closely as possible under the participant budget.Evaluation network · per MLPWidth nnn256256256Hidden layers LLL323232Weight initializationHe-Gaussian · variance 2/n2/n2/nInput distributionX∼N(0,In)X \sim \mathcal{N}(0, I_n)X∼N(0,In​)Prediction shape32×25632 \times 25632×256 matrixPrimary metricFinal-layer MSE vs. a high-budget Monte Carlo referenceImportantThe submission is executable code, not a prediction file. The grader runs your estimator against held-out MLPs and scores the returned activation matrix.03Compute model and constraintsThe competition is budgeted by analytical FLOPs, not by who owns the fastest machine.The accounting library is flopscope, a NumPy-compatible interface that counts floating-point operations for instrumented operations. Code written through flopscope.numpy is charged analytically. Uninstrumented computation is allowed, but residual wall-clock time is converted back into FLOPs at an unfavorable rate.Cm=Fm+λRmC_m = F_m + \lambda R_mCm​=Fm​+λRm​3Here FmF_mFm​ is the analytical FLOP count for MLP mmm, RmR_mRm​ is residual wall-clock time outside instrumented operations, and λ\lambdaλ is the conversion rate published in the starter kit.import flopscope as flops import flopscope.numpy as fnp def predict(mlp, budget): mus = [] mu = fnp.zeros(mlp.width) var = fnp.ones(mlp.width) for w in mlp.weights: mu_pre = w.T @ mu var_pre = (w * w).T @ var sigma_pre = fnp.sqrt(fnp.maximum(var_pre, 1e-12)) alpha = mu_pre / sigma_pre mu = mu_pre * flops.stats.norm.cdf(alpha) + sigma_pre * flops.stats.norm.pdf(alpha) mus.append(mu) return fnp.stack(mus)Budget ruleStay within the per-MLP budget on every network. Over-budget runs, exceptions, invalid shapes, non-finite values, memory failures, or wall-clock guard failures receive the zero-prediction fallback for that MLP.Grader environmentYour code runs on one physical core (2 vCPUs); the flopscope backend runs on seven physical cores (14 vCPUs). Submissions are CPU-only with 64 GB RAM, disabled network access, and a 60-second hard wall-clock cap per MLP.Fig. 2Bₘ — per-MLP budget (2.72 × 10¹¹ FLOPs)10⁷10⁹10¹²10¹⁵10⁰10⁻³10⁻⁶10⁻⁹10⁻¹²FLOPs per MLP (log scale) →Mean propagation9.5 × 10⁻⁴Covariance propagation8.4 × 10⁻⁵source · 100 MLPs · ARC Phase 1 dataBlack-box baselineMonte Carloconvergence.Monte Carlo on 100 randomMLPs. Red bands showvariation across networks;the dashed line is themean MSE (the scored bar).White-box points sit belowthe dashed line — lower errorthan sampling at equal compute.Figure 2Monte Carlo convergence. Pure Monte Carlo estimates the final-layer activation mean by buying more forward passes; plotted against the compute spent, its mean squared error falls steadily as the per-MLP FLOPs budget grows.The red bands summarize Monte Carlo error across 100 random MLPs as the sampling budget — the compute spent on forward-pass sampling per MLP — increases along the horizontal axis, measured in FLOPs (one black-box forward pass ≈4.24×106\approx 4.24 \times 10^{6}≈4.24×106 FLOPs). The dashed line is the mean final-layer MSE across MLPs — the quantity scored against (E[MSE] = σ²/N), so the Monte Carlo @ Bₘ reference lands on it; nested bands show between-MLP spread (median and percentiles). The vertical marker is the per-MLP budget Bm≈2.72×1011B_m \approx 2.72 \times 10^{11}Bm​≈2.72×1011 FLOPs. Baseline white-box methods such as mean propagation and covariance propagation appear as points because they spend compute inspecting weights and propagating distributional statistics rather than only sampling inputs. The challenge is to move below the red convergence curve under the same effective-compute budget: lower final-layer MSE, without exceeding Cm≤BmC_m \le B_mCm​≤Bm​.04Evaluation and scoringThe live leaderboard is useful feedback; the final ranking comes from a fresh private rerun.For each evaluation MLP, the grader computes the final-layer mean squared error between your prediction and the Monte Carlo reference:MSEfinal,m=1n∑i(Y^L,i−YL,i)2\mathrm{MSE}_{\mathrm{final},m} = \frac{1}{n}\sum_i \left(\hat{Y}_{L,i} - Y_{L,i}\right)^2MSEfinal,m​=n1​i∑​(Y^L,i​−YL,i​)24The per-MLP leaderboard score multiplies this by a compute-usage factor. Staying under the budget can help, but the improvement is capped so that an extremely cheap but inaccurate estimator cannot dominate by spending little compute.sm=MSEfinal,m⋅max⁡(0.1,  Cm/Bm)s_m = \mathrm{MSE}_{\mathrm{final},m} \cdot \max\left(0.1,\; C_m / B_m\right)sm​=MSEfinal,m​⋅max(0.1,Cm​/Bm​)5The overall leaderboard score is the average of sms_msm​ across the evaluation suite. Lower is better.All-layer MSE, averaged across all L×nL \times nL×n hidden activations, is reported as a secondary diagnostic. It helps reveal where approximation error accumulates across layers, but the primary score is the final-layer score.During each official phase, the grader evaluates submissions on a private suite of 100 randomly generated MLPs. Fifty contribute to live public feedback, while fifty are withheld until the phase closes. This keeps the leaderboard informative without making it too easy to overfit to visible scores.After each phase closes, each team's nominated submissions — up to two per phase, or your two highest-ranked public submissions if you nominate none — are rerun on a separate, freshly generated private test suite of new MLPs from the same distribution, using private seeds held out from both phases. Prize ranking is decided exclusively from this private rerun, not from the best public-leaderboard score observed during the competition.If leading submissions are statistically indistinguishable after the Private Re-evaluation, the Rules allow the Sponsor to generate additional MLPs from the same distribution for statistical disambiguation. If submissions remain tied after that, the tied ranks share the combined prize amounts for those positions.Public score vs. final prize rankThe public board helps you iterate. The final private rerun decides prize ranking.Failed-run fallbackIf a submission exceeds the budget, raises an exception, returns invalid shapes or non-finite values, exhausts memory, or trips an operational guard on a given MLP, the grader substitutes a zero prediction for that MLP and continues. No compute discount is applied to the fallback.Do not overfit the public boardThe final private suite uses different random MLPs. Strong submissions should generalize across the published generative distribution, not exploit visible leaderboard instances.05How to participateStart locally, validate the estimator contract, then submit a packaged tarball through AIcrowd.git clone https://github.com/AIcrowd/whest-starterkit.git cd whest-starterkit uv sync uv run python estimator.pyThe starter kit is structured as a staged ladder. Point your local runs at the public dataset's Mini split while you iterate. A good first milestone is one valid end-to-end submission, even before you have a strong score.1Iterate locallyuv run python estimator.pyCheck the math against a local Monte Carlo harness.2Validate contractuv run whest validate --estimator estimator.pyCatch shape, type, and packaging issues early.3Run on the public setuv run whest run --estimator estimator.py \ --dataset hf://aicrowd/arc-whestbench-public-2026@v1-phase1 \ --split mini --runner localReal scoring against the public Mini split in a debuggable process.4Subprocess runneruv run whest run --estimator estimator.py \ --dataset hf://aicrowd/arc-whestbench-public-2026@v1-phase1 \ --split mini --runner subprocessTest isolation closer to the grader.5Package and submituv run whest package -o submission.tar.gzuv run whest loginuv run whest submit submission.tar.gzBuild the tarball, authenticate, and upload your submission to AIcrowd.First milestoneYour first goal should be one valid end-to-end submission. Once the contract, packaging, and grader path work, you can improve the estimator.Open starter kitMake a submission06Rules that matter for first submissionThis is not a substitute for the Rules page, but it covers the constraints most likely to affect your first estimator.Submission formatSubmit executable code, including an estimator.py that follows the starter-kit contract. Do not submit prediction files.Submission capEach team may submit up to 50 entries per UTC day during each official Phase. The UTC-day counter resets at 00:00 UTC.TeamsUp to five eligible individuals per team, finalized by July 31 September 5, 2026, 23:59 UTC.HardwareCPU-only. Your code runs on one physical core (2 vCPUs); the flopscope backend runs on seven physical cores (14 vCPUs). 64 GB RAM, disabled network, and a 60-second hard wall-clock cap per MLP.Network accessNetwork access is disabled during evaluation. Bundle any allowed weights, lookup tables, dependencies, or precomputed artifacts in the submission tarball.Do not tamperDo not modify flopscope, read private seeds, access grader internals, or otherwise circumvent budget enforcement.LLM & autoresearchLLM-assisted and agentic development is welcome. You remain responsible for compliance, attribution, reproducibility, and any technical-writeup disclosures required by the Rules.Final submissionFor each phase, nominate up to two valid submissions for that phase's private rerun (Phase 2 by Sep 19, 2026, 23:59 UTC). With no nomination, Sponsor uses your two highest-ranked valid submissions on that phase's public leaderboard.Prize rankingDecided exclusively by the final private leaderboard from the fresh Private Re-evaluation suite. The public leaderboard is for iteration and does not determine prizes.Rules governIf anything on this Overview page conflicts with the official Rules or current starter kit, follow the Rules and starter kit.Autoresearch is welcomeUse LLMs, code agents, public resources, and metric-driven iteration if they help you discover better estimators.The one boundary is rule evasion: don't automate registration or mass uploads, tamper with flopscope, read private grader materials, or submit work you can't verify.Read the official policy →07Prizes and recognitionWhestBench rewards both leaderboard performance and algorithmic contribution.At launch, WhestBench has USD 150,000+ in prizes and recognition planned across two official phases. The current Rules specify $150,000 USD in total place-prize ARV — $50,000 in Phase 1 and $100,000 in Phase 2 — split across score-based and algorithmic contribution prizes. Sponsor may increase prize amounts or offer additional prizes, and any changes will be announced on the Competition Site.Total prize poolAcross two official phases · USD ARVCombined$150,000+By place & phasePhase 1Phase 21st place$25,000$50,0002nd place$10,000$20,0003rd place$5,000$10,000Algorithmic contributionBest technical contribution to mechanistic estimation — judged on score, algorithmic ideas, and write-up quality.$10,000$20,000Subtotal$50,000$100,000Beyond rankCommunity contributionDiscretionary recognition for helpful competition contributions — awarded per contributor.$500–5,000All amounts in USD · ARV.Algorithmic contributionHow to submit — PDF write-up + submission ID, deadlines — and how ARC judges these prizesHow to submitAn algorithmic contribution prize entry consists of:A PDF technical write-up describing your approach, the core ideas behind it, and the evidence/results supporting it (negative results and ablations are welcome), andExactly one submission ID of a submission that was successfully evaluated by the grader of that phase. The write-up must describe the approach behind that specific submission. No separate code upload is needed — your submission ID already points to your validated code package on our servers.You can submit in either of two ways:Privately, by email to arc-whestbench@aicrowd.com, orPublicly, as a post on the Challenge Discussion Forum. Public write-ups will additionally be considered for Community Contribution Prizes, where applicable.One write-up · one submission IDIn both cases, clearly state the submission ID your write-up refers to. Each write-up maps to exactly one submission that was successfully graded in that phase.Deadlines — you get one extra week after each phase closes to finish the write-up:Aug 7 Aug 17 · 23:59Phase 1 algorithmic-contribution write-up deadlineSep 26 · 23:59Phase 2 algorithmic-contribution write-up deadlineAll times are UTC.The extra week is for writing onlyOnce a phase ends, its evaluator closes — you will not be able to make new submissions or re-grade anything for that phase. The submission you reference must already be successfully graded before the phase deadline (July 31 August 10 for Phase 1; September 19 for Phase 2).How ARC judgesGuidance from the Alignment Research Center (ARC)These prizes will be awarded at ARC's discretion to the method we think most improves our understanding of white-box estimation for random MLPs. We are most interested in "mechanistic" estimation methods, as discussed in our blog posts on competing with sampling and mechanistic estimation for wide random MLPs. We are less interested in methods that rely heavily on sampling, fine-tuned constants, careful performance optimization, and opaque LLM-optimized code (although clever sampling-based methods are of interest if they rely on interesting structural observations, and LLM-written code is of interest providing it can be deciphered).Technical writeups. The chance of a submission receiving an algorithmic contribution prize is greatly increased by the inclusion of a technical writeup explaining the algorithmic approach used and how it was developed. We will likely start by reading the technical writeup for the highest-scoring submissions, and award the prize to the submission where novel "mechanistic" ideas made the largest improvement to performance over previously-known methods.LLM usage. We are ultimately interested in the quality of the algorithmic contribution itself, regardless of how it was obtained, and LLM usage is encouraged. However, contestants should be fully transparent about the extent to which LLMs were used to generate code and/or portions of technical writeups. If contestants have significant uncertainty about how and why their code actually works, the relevant portions of the technical writeup should be appropriately hedged and/or labeled as guesswork (for example, "the LLM gave this explanation, which we did not validate/which we validated by ..."). If we notice unhedged, dubious claims, then we are likely to be more skeptical about the remaining content and may skip over submissions entirely.This guidance is also posted as ARC's guidelines post on the forum.Recognition beyond rankStrong submissions may be valuable even when they are not first on the leaderboard. Clear explanations, useful algorithmic ideas, helpful bug reports, and community contributions may be recognized according to the Rules and any later announcements on the Competition Site.Winning place-prize submissions are subject to verification and open-source release requirements described in the Rules. The current Rules require place-prize winners to release the prize-determining solution code and required artifacts under an OSI-approved open-source license within 30 days of winner notification, and to keep the release publicly accessible for at least three years.08TimelineNowMay 28Jun 18Aug 12Sep 19Oct 1Warm-upsubmissions openPhase 1public boardPhase 2final submissionsDeadlinesubmissions closeWinnersTimelineFrom warm-up to tentative winners. The strip gives the shape of the schedule at a glance; the table below is the source for exact dates and times.Warm-upMay 28 – Jun 17May 28 · 00:00Resources released; submissions openJun 17 · 23:59Warm-up round endsPhase 1 · public leaderboardJun 18 – Jul 31 Aug 10Jun 18 · 00:00Public leaderboard opensJul 31 Aug 10 · 23:59Phase 1 ends; submissions closeAug 11–17Phase 1 private re-evaluation on a fresh private suite (up to 2 nominated submissions per team)Aug 7 Aug 17 · 23:59Phase 1 algorithmic-contribution write-up deadline (write-up only; the Phase 1 evaluator is closed)Phase 2 · final submissionsAug 1 Aug 12 – Sep 19Aug 1 Aug 12 · 00:00Final submission period opensSep 5 · 23:59Registration and team freezeSep 19 · 23:59Phase 2 ends; submissions closeSep 26 · 23:59Phase 2 algorithmic-contribution write-up deadline (write-up only; the Phase 2 evaluator is closed)Evaluation & resultsSep 20 – Oct 1Sep 20–30Private re-evaluation on a fresh held-out suite≈ Oct 1Tentative winner announcementAll times are UTC.09Resources and contactUse the starter kit for implementation details, the Rules page for official constraints, and the forum for public questions.CompeteParticipate on AIcrowdRegister, form a team, and submit through the platform.Challenge RulesOfficial constraints, eligibility, and prize terms.BuildWhestBench starter kitClone, implement your estimator, validate, and package a submission.Public dataset1,100 random MLPs on Hugging Face · Mini and Full splits.flopscopeThe NumPy-compatible FLOP-accounting library the grader uses.WhestBench ExplorerInspect generated MLPs and their activation statistics.Community & supportDiscussion forumPublic questions, clarifications, and announcements.GitHub IssuesReport bugs in the starter kit or flopscope.arc-whestbench@aicrowd.comPrivate or administrative matters.ResearchCompanion paperThe research behind the benchmark — arXiv:2605.05179.ARC announcementThe Alignment Research Center research post.If you use WhestBench in academic work, cite the companion paper:Wilson Wu, Victor Lecomte, Michael Winer, George Robinson, Jacob Hilton, and Paul Christiano. "Estimating the expected output of wide random MLPs more efficiently than sampling." arXiv:2605.05179, 2026.The challenge is organized by Alignment Research Center in partnership with AIcrowd.Ready to submit your first estimator?Start with the starter kit, validate locally, then submit through AIcrowd. The first useful milestone is not leaderboard rank—it is one valid end-to-end submission.Make a submissionOpen starter kitRead the Rules

#llm#Overview#Why random
Prize
$150,000
Details
🤖
aicrowd
ACTIVE|Lvl 5/10

ARC White-Box Estimation Challenge 2026

00Overview01Why random02The task03Compute04Evaluation05Participate06Rules07Prizes08Timeline09ResourcesSubmissions openStarter kitWhestBenchflopscopeHF datasetWhen can we know what a neural network does without running it?The ARC White-Box Estimation Challenge is a contest in compute-efficient mechanistic estimation. Given the weights of a neural network, can you predict its expected per-neuron activations more accurately than running it many times?The obvious way to learn how a model behaves is to run it many times and average what you observe. That works well when the behavior is common, cheap to elicit, and easy to sample. But when the behavior is rare, high-variance, or unlikely to appear in obvious test cases, brute-force testing can become an expensive way to learn very little.The ARC White-Box Estimation Challenge turns that question into a controlled benchmark. Participants receive randomly initialized ReLU MLPs and build executable estimators that predict each neuron's expected post-ReLU activation under standard-normal inputs.The goal is simple to state: beat comparable black-box sampling under a shared compute budget by using the network's weights. The strongest submissions may be Monte Carlo, white-box, hybrid, LLM-assisted, or something unexpected—the leaderboard will decide.TaskExecutable estimatorInputWeights + budgetOutputExpected activationsMetricFinal-layer MSELatestUpdateAug 3Phase 1 update: flopscope v0.10.0, cost-model fixes, residual-time safeguards, and updated deadlines↗AnnouncementJun 18Phase 1 launched — deeper models, and increased prizes.↗All updates on the forum↗Official factsPrize pool$150,000 USD ARVTwo phases · $50K Phase 1 + $100K Phase 2 · places + algorithmicSubmissions openMay 28, 2026 · 00:00 UTCPhase 2 closesSep 19, 2026 · 23:59 UTCDaily limit50 entries per team · per UTC day, each phaseGraderCPU-only 1 core (2 vCPU) for your code · 7-core flopscope backend · 64 GB · no networkHard cap60 s per MLPFinal rankingFresh private rerun of each team's up to two nominated submissions, per phaseFig. 1 < !-- Header -- > CHALLENGE · ESTIMATE THE EXPECTATION .katex-display{margin:0 !important;} Y^L,j≈EX∼N(0,In) ⁣[hj(L)(X)]\hat Y_{L,j} \approx \mathbb{E}_{X\sim\mathcal{N}(0,I_n)}\!\left[h^{(L)}_{j}(X)\right]Y^L,j​≈EX∼N(0,In​)​[hj(L)​(X)] < !-- Column labels -- > INPUT NETWORK · RANDOM ReLU MLP OUTPUT · Ê[ h⁽ᶽ⁾₃(X) ] < !-- Input PDF -- > +2 0 −2 xᵢ +0.71 < !-- Network -- > h⁽ᶽ⁾₃ < !-- Output axis -- > 7.5 8.0 8.5 9.0 9.5 h⁽ᶽ⁾₃(X) · ×10⁻³ Ê 8.6 µ̂ 8.3 Δ 3×10⁻⁴ERROR≈3.5% rel < !-- Method legend -- > MONTE CARLOBLACK-BOX · SAMPLE & AVERAGE SAMPLING… ANALYTICALWHITE-BOX · PROPAGATE THE DISTRIBUTION PROPAGATING… < !-- Cost band -- > ← the contest lives here ≈ 15,000× YOUR BUDGET 2.72×10¹¹ FLOPs / MLP MONTE CARLO REFERENCE 4.24×10¹⁵ FLOPs / MLP 10¹¹ 10¹² 10¹³ 10¹⁴ 10¹⁵ 10¹⁶ FLOPs (log₁₀ scale) < !-- The question -- > Can you beat sampling? < !-- Replay -- > REPLAY Figure 1The estimation problem, as a distributional computation. A generated ReLU MLP receives Gaussian inputs X∼N(0,In)X \sim \mathcal{N}(0, I_n)X∼N(0,In​), applies h(ℓ)=ReLU ⁣(W(ℓ)h(ℓ−1))h^{(\ell)} = \mathrm{ReLU}\!\left(W^{(\ell)}h^{(\ell-1)}\right)h(ℓ)=ReLU(W(ℓ)h(ℓ−1)), and the submission must estimate E ⁣[hi(ℓ)(X)]\mathbb{E}\!\left[h^{(\ell)}_i(X)\right]E[hi(ℓ)​(X)] for every hidden-layer neuron.Read the animation as two ways to estimate the same activation-mean matrix. The black-box path samples inputs, runs the network, and averages observed activations until the Monte Carlo mean stabilizes. The white-box path inspects W(1),…,W(L)W^{(1)}, \dots, W^{(L)}W(1),…,W(L) and propagates enough distributional information to predict the same means under the participant budget. The target is an organizers' high-budget Monte Carlo reference of approximately 4.24×10154.24 \times 10^{15}4.24×1015 FLOPs, compared with a participant budget of approximately 2.72×10112.72 \times 10^{11}2.72×1011 FLOPs per MLP—roughly a 15,000×15{,}000\times15,000× compute gap.Prizes$150K+Cash PrizesPrediction shape32 × 256hidden activationsBudget / MLP2.72e11FLOPsPhase 2 closesSep 192026 · 23:59 UTC01Why this starts with random networksThe benchmark isolates one hard part of white-box estimation: tracking how distributions move through nonlinear layers.White-box estimation for trained networks is the destination, not the starting line. Trained models introduce many confounders at once: data, optimization, learned structure, task semantics, and evaluation ambiguity. WhestBench begins with randomly initialized networks so participants can focus on the estimation problem in a simplified setting.The networks are synthetic, but the question is real. Given access to the weights, can an algorithm reason about the distribution of hidden activations more efficiently than repeatedly sampling inputs and averaging outputs?Random ReLU MLPs retain the same basic problem structure: each layer transforms a distribution, the ReLU nonlinearity reshapes it, and approximation error can accumulate with depth. The first challenge is to develop methods that work in this controlled setting; later work can ask how those methods adapt as networks acquire structure during training.The benchmark is controlled, but not trivial: the expected activation has no closed form for the full network, and sampling improves only slowly with more compute.Why not trained models first?Random networks offer a simplified setting for compute-efficient estimation while being an important stepping stone towards trained models.02The taskFor each MLP, return a matrix of expected post-ReLU activation means.For each evaluation network MθM_\thetaMθ​, your estimator receives the MLP weights and a compute budget. It must return an L×nL \times nL×n matrix Y^\hat{Y}Y^. Entry (ℓ,i)(\ell, i)(ℓ,i) should estimate the expected post-ReLU activation of neuron iii in hidden layer ℓ\ellℓ when inputs are drawn from a standard Gaussian distribution.h(0)=X,h(ℓ)=ReLU(W(ℓ)h(ℓ−1)),ℓ=1,…,Lh^{(0)} = X, \qquad h^{(\ell)} = \mathrm{ReLU}\left(W^{(\ell)}h^{(\ell-1)}\right), \quad \ell = 1, \dots, Lh(0)=X,h(ℓ)=ReLU(W(ℓ)h(ℓ−1)),ℓ=1,…,L1Y^ℓ,i≈EX∼N(0,In)[hi(ℓ)(X)]\hat{Y}_{\ell,i} \approx \mathbb{E}_{X \sim \mathcal{N}(0, I_n)}\left[h^{(\ell)}_i(X)\right]Y^ℓ,i​≈EX∼N(0,In​)​[hi(ℓ)​(X)]2The reference target is estimated by the organizers with a much larger Monte Carlo budget than participants receive. Your job is to match that reference as closely as possible under the participant budget.Evaluation network · per MLPWidth nnn256256256Hidden layers LLL323232Weight initializationHe-Gaussian · variance 2/n2/n2/nInput distributionX∼N(0,In)X \sim \mathcal{N}(0, I_n)X∼N(0,In​)Prediction shape32×25632 \times 25632×256 matrixPrimary metricFinal-layer MSE vs. a high-budget Monte Carlo referenceImportantThe submission is executable code, not a prediction file. The grader runs your estimator against held-out MLPs and scores the returned activation matrix.03Compute model and constraintsThe competition is budgeted by analytical FLOPs, not by who owns the fastest machine.The accounting library is flopscope, a NumPy-compatible interface that counts floating-point operations for instrumented operations. Code written through flopscope.numpy is charged analytically. Uninstrumented computation is allowed, but residual wall-clock time is converted back into FLOPs at an unfavorable rate.Cm=Fm+λRmC_m = F_m + \lambda R_mCm​=Fm​+λRm​3Here FmF_mFm​ is the analytical FLOP count for MLP mmm, RmR_mRm​ is residual wall-clock time outside instrumented operations, and λ\lambdaλ is the conversion rate published in the starter kit.import flopscope as flops import flopscope.numpy as fnp def predict(mlp, budget): mus = [] mu = fnp.zeros(mlp.width) var = fnp.ones(mlp.width) for w in mlp.weights: mu_pre = w.T @ mu var_pre = (w * w).T @ var sigma_pre = fnp.sqrt(fnp.maximum(var_pre, 1e-12)) alpha = mu_pre / sigma_pre mu = mu_pre * flops.stats.norm.cdf(alpha) + sigma_pre * flops.stats.norm.pdf(alpha) mus.append(mu) return fnp.stack(mus)Budget ruleStay within the per-MLP budget on every network. Over-budget runs, exceptions, invalid shapes, non-finite values, memory failures, or wall-clock guard failures receive the zero-prediction fallback for that MLP.Grader environmentYour code runs on one physical core (2 vCPUs); the flopscope backend runs on seven physical cores (14 vCPUs). Submissions are CPU-only with 64 GB RAM, disabled network access, and a 60-second hard wall-clock cap per MLP.Fig. 2Bₘ — per-MLP budget (2.72 × 10¹¹ FLOPs)10⁷10⁹10¹²10¹⁵10⁰10⁻³10⁻⁶10⁻⁹10⁻¹²FLOPs per MLP (log scale) →Mean propagation9.5 × 10⁻⁴Covariance propagation8.4 × 10⁻⁵source · 100 MLPs · ARC Phase 1 dataBlack-box baselineMonte Carloconvergence.Monte Carlo on 100 randomMLPs. Red bands showvariation across networks;the dashed line is themean MSE (the scored bar).White-box points sit belowthe dashed line — lower errorthan sampling at equal compute.Figure 2Monte Carlo convergence. Pure Monte Carlo estimates the final-layer activation mean by buying more forward passes; plotted against the compute spent, its mean squared error falls steadily as the per-MLP FLOPs budget grows.The red bands summarize Monte Carlo error across 100 random MLPs as the sampling budget — the compute spent on forward-pass sampling per MLP — increases along the horizontal axis, measured in FLOPs (one black-box forward pass ≈4.24×106\approx 4.24 \times 10^{6}≈4.24×106 FLOPs). The dashed line is the mean final-layer MSE across MLPs — the quantity scored against (E[MSE] = σ²/N), so the Monte Carlo @ Bₘ reference lands on it; nested bands show between-MLP spread (median and percentiles). The vertical marker is the per-MLP budget Bm≈2.72×1011B_m \approx 2.72 \times 10^{11}Bm​≈2.72×1011 FLOPs. Baseline white-box methods such as mean propagation and covariance propagation appear as points because they spend compute inspecting weights and propagating distributional statistics rather than only sampling inputs. The challenge is to move below the red convergence curve under the same effective-compute budget: lower final-layer MSE, without exceeding Cm≤BmC_m \le B_mCm​≤Bm​.04Evaluation and scoringThe live leaderboard is useful feedback; the final ranking comes from a fresh private rerun.For each evaluation MLP, the grader computes the final-layer mean squared error between your prediction and the Monte Carlo reference:MSEfinal,m=1n∑i(Y^L,i−YL,i)2\mathrm{MSE}_{\mathrm{final},m} = \frac{1}{n}\sum_i \left(\hat{Y}_{L,i} - Y_{L,i}\right)^2MSEfinal,m​=n1​i∑​(Y^L,i​−YL,i​)24The per-MLP leaderboard score multiplies this by a compute-usage factor. Staying under the budget can help, but the improvement is capped so that an extremely cheap but inaccurate estimator cannot dominate by spending little compute.sm=MSEfinal,m⋅max⁡(0.1,  Cm/Bm)s_m = \mathrm{MSE}_{\mathrm{final},m} \cdot \max\left(0.1,\; C_m / B_m\right)sm​=MSEfinal,m​⋅max(0.1,Cm​/Bm​)5The overall leaderboard score is the average of sms_msm​ across the evaluation suite. Lower is better.All-layer MSE, averaged across all L×nL \times nL×n hidden activations, is reported as a secondary diagnostic. It helps reveal where approximation error accumulates across layers, but the primary score is the final-layer score.During each official phase, the grader evaluates submissions on a private suite of 100 randomly generated MLPs. Fifty contribute to live public feedback, while fifty are withheld until the phase closes. This keeps the leaderboard informative without making it too easy to overfit to visible scores.After each phase closes, each team's nominated submissions — up to two per phase, or your two highest-ranked public submissions if you nominate none — are rerun on a separate, freshly generated private test suite of new MLPs from the same distribution, using private seeds held out from both phases. Prize ranking is decided exclusively from this private rerun, not from the best public-leaderboard score observed during the competition.If leading submissions are statistically indistinguishable after the Private Re-evaluation, the Rules allow the Sponsor to generate additional MLPs from the same distribution for statistical disambiguation. If submissions remain tied after that, the tied ranks share the combined prize amounts for those positions.Public score vs. final prize rankThe public board helps you iterate. The final private rerun decides prize ranking.Failed-run fallbackIf a submission exceeds the budget, raises an exception, returns invalid shapes or non-finite values, exhausts memory, or trips an operational guard on a given MLP, the grader substitutes a zero prediction for that MLP and continues. No compute discount is applied to the fallback.Do not overfit the public boardThe final private suite uses different random MLPs. Strong submissions should generalize across the published generative distribution, not exploit visible leaderboard instances.05How to participateStart locally, validate the estimator contract, then submit a packaged tarball through AIcrowd.git clone https://github.com/AIcrowd/whest-starterkit.git cd whest-starterkit uv sync uv run python estimator.pyThe starter kit is structured as a staged ladder. Point your local runs at the public dataset's Mini split while you iterate. A good first milestone is one valid end-to-end submission, even before you have a strong score.1Iterate locallyuv run python estimator.pyCheck the math against a local Monte Carlo harness.2Validate contractuv run whest validate --estimator estimator.pyCatch shape, type, and packaging issues early.3Run on the public setuv run whest run --estimator estimator.py \ --dataset hf://aicrowd/arc-whestbench-public-2026@v1-phase1 \ --split mini --runner localReal scoring against the public Mini split in a debuggable process.4Subprocess runneruv run whest run --estimator estimator.py \ --dataset hf://aicrowd/arc-whestbench-public-2026@v1-phase1 \ --split mini --runner subprocessTest isolation closer to the grader.5Package and submituv run whest package -o submission.tar.gzuv run whest loginuv run whest submit submission.tar.gzBuild the tarball, authenticate, and upload your submission to AIcrowd.First milestoneYour first goal should be one valid end-to-end submission. Once the contract, packaging, and grader path work, you can improve the estimator.Open starter kitMake a submission06Rules that matter for first submissionThis is not a substitute for the Rules page, but it covers the constraints most likely to affect your first estimator.Submission formatSubmit executable code, including an estimator.py that follows the starter-kit contract. Do not submit prediction files.Submission capEach team may submit up to 50 entries per UTC day during each official Phase. The UTC-day counter resets at 00:00 UTC.TeamsUp to five eligible individuals per team, finalized by July 31 September 5, 2026, 23:59 UTC.HardwareCPU-only. Your code runs on one physical core (2 vCPUs); the flopscope backend runs on seven physical cores (14 vCPUs). 64 GB RAM, disabled network, and a 60-second hard wall-clock cap per MLP.Network accessNetwork access is disabled during evaluation. Bundle any allowed weights, lookup tables, dependencies, or precomputed artifacts in the submission tarball.Do not tamperDo not modify flopscope, read private seeds, access grader internals, or otherwise circumvent budget enforcement.LLM & autoresearchLLM-assisted and agentic development is welcome. You remain responsible for compliance, attribution, reproducibility, and any technical-writeup disclosures required by the Rules.Final submissionFor each phase, nominate up to two valid submissions for that phase's private rerun (Phase 2 by Sep 19, 2026, 23:59 UTC). With no nomination, Sponsor uses your two highest-ranked valid submissions on that phase's public leaderboard.Prize rankingDecided exclusively by the final private leaderboard from the fresh Private Re-evaluation suite. The public leaderboard is for iteration and does not determine prizes.Rules governIf anything on this Overview page conflicts with the official Rules or current starter kit, follow the Rules and starter kit.Autoresearch is welcomeUse LLMs, code agents, public resources, and metric-driven iteration if they help you discover better estimators.The one boundary is rule evasion: don't automate registration or mass uploads, tamper with flopscope, read private grader materials, or submit work you can't verify.Read the official policy →07Prizes and recognitionWhestBench rewards both leaderboard performance and algorithmic contribution.At launch, WhestBench has USD 150,000+ in prizes and recognition planned across two official phases. The current Rules specify $150,000 USD in total place-prize ARV — $50,000 in Phase 1 and $100,000 in Phase 2 — split across score-based and algorithmic contribution prizes. Sponsor may increase prize amounts or offer additional prizes, and any changes will be announced on the Competition Site.Total prize poolAcross two official phases · USD ARVCombined$150,000+By place & phasePhase 1Phase 21st place$25,000$50,0002nd place$10,000$20,0003rd place$5,000$10,000Algorithmic contributionBest technical contribution to mechanistic estimation — judged on score, algorithmic ideas, and write-up quality.$10,000$20,000Subtotal$50,000$100,000Beyond rankCommunity contributionDiscretionary recognition for helpful competition contributions — awarded per contributor.$500–5,000All amounts in USD · ARV.Algorithmic contributionHow to submit — PDF write-up + submission ID, deadlines — and how ARC judges these prizesHow to submitAn algorithmic contribution prize entry consists of:A PDF technical write-up describing your approach, the core ideas behind it, and the evidence/results supporting it (negative results and ablations are welcome), andExactly one submission ID of a submission that was successfully evaluated by the grader of that phase. The write-up must describe the approach behind that specific submission. No separate code upload is needed — your submission ID already points to your validated code package on our servers.You can submit in either of two ways:Privately, by email to arc-whestbench@aicrowd.com, orPublicly, as a post on the Challenge Discussion Forum. Public write-ups will additionally be considered for Community Contribution Prizes, where applicable.One write-up · one submission IDIn both cases, clearly state the submission ID your write-up refers to. Each write-up maps to exactly one submission that was successfully graded in that phase.Deadlines — you get one extra week after each phase closes to finish the write-up:Aug 7 Aug 17 · 23:59Phase 1 algorithmic-contribution write-up deadlineSep 26 · 23:59Phase 2 algorithmic-contribution write-up deadlineAll times are UTC.The extra week is for writing onlyOnce a phase ends, its evaluator closes — you will not be able to make new submissions or re-grade anything for that phase. The submission you reference must already be successfully graded before the phase deadline (July 31 August 10 for Phase 1; September 19 for Phase 2).How ARC judgesGuidance from the Alignment Research Center (ARC)These prizes will be awarded at ARC's discretion to the method we think most improves our understanding of white-box estimation for random MLPs. We are most interested in "mechanistic" estimation methods, as discussed in our blog posts on competing with sampling and mechanistic estimation for wide random MLPs. We are less interested in methods that rely heavily on sampling, fine-tuned constants, careful performance optimization, and opaque LLM-optimized code (although clever sampling-based methods are of interest if they rely on interesting structural observations, and LLM-written code is of interest providing it can be deciphered).Technical writeups. The chance of a submission receiving an algorithmic contribution prize is greatly increased by the inclusion of a technical writeup explaining the algorithmic approach used and how it was developed. We will likely start by reading the technical writeup for the highest-scoring submissions, and award the prize to the submission where novel "mechanistic" ideas made the largest improvement to performance over previously-known methods.LLM usage. We are ultimately interested in the quality of the algorithmic contribution itself, regardless of how it was obtained, and LLM usage is encouraged. However, contestants should be fully transparent about the extent to which LLMs were used to generate code and/or portions of technical writeups. If contestants have significant uncertainty about how and why their code actually works, the relevant portions of the technical writeup should be appropriately hedged and/or labeled as guesswork (for example, "the LLM gave this explanation, which we did not validate/which we validated by ..."). If we notice unhedged, dubious claims, then we are likely to be more skeptical about the remaining content and may skip over submissions entirely.This guidance is also posted as ARC's guidelines post on the forum.Recognition beyond rankStrong submissions may be valuable even when they are not first on the leaderboard. Clear explanations, useful algorithmic ideas, helpful bug reports, and community contributions may be recognized according to the Rules and any later announcements on the Competition Site.Winning place-prize submissions are subject to verification and open-source release requirements described in the Rules. The current Rules require place-prize winners to release the prize-determining solution code and required artifacts under an OSI-approved open-source license within 30 days of winner notification, and to keep the release publicly accessible for at least three years.08TimelineNowMay 28Jun 18Aug 12Sep 19Oct 1Warm-upsubmissions openPhase 1public boardPhase 2final submissionsDeadlinesubmissions closeWinnersTimelineFrom warm-up to tentative winners. The strip gives the shape of the schedule at a glance; the table below is the source for exact dates and times.Warm-upMay 28 – Jun 17May 28 · 00:00Resources released; submissions openJun 17 · 23:59Warm-up round endsPhase 1 · public leaderboardJun 18 – Jul 31 Aug 10Jun 18 · 00:00Public leaderboard opensJul 31 Aug 10 · 23:59Phase 1 ends; submissions closeAug 11–17Phase 1 private re-evaluation on a fresh private suite (up to 2 nominated submissions per team)Aug 7 Aug 17 · 23:59Phase 1 algorithmic-contribution write-up deadline (write-up only; the Phase 1 evaluator is closed)Phase 2 · final submissionsAug 1 Aug 12 – Sep 19Aug 1 Aug 12 · 00:00Final submission period opensSep 5 · 23:59Registration and team freezeSep 19 · 23:59Phase 2 ends; submissions closeSep 26 · 23:59Phase 2 algorithmic-contribution write-up deadline (write-up only; the Phase 2 evaluator is closed)Evaluation & resultsSep 20 – Oct 1Sep 20–30Private re-evaluation on a fresh held-out suite≈ Oct 1Tentative winner announcementAll times are UTC.09Resources and contactUse the starter kit for implementation details, the Rules page for official constraints, and the forum for public questions.CompeteParticipate on AIcrowdRegister, form a team, and submit through the platform.Challenge RulesOfficial constraints, eligibility, and prize terms.BuildWhestBench starter kitClone, implement your estimator, validate, and package a submission.Public dataset1,100 random MLPs on Hugging Face · Mini and Full splits.flopscopeThe NumPy-compatible FLOP-accounting library the grader uses.WhestBench ExplorerInspect generated MLPs and their activation statistics.Community & supportDiscussion forumPublic questions, clarifications, and announcements.GitHub IssuesReport bugs in the starter kit or flopscope.arc-whestbench@aicrowd.comPrivate or administrative matters.ResearchCompanion paperThe research behind the benchmark — arXiv:2605.05179.ARC announcementThe Alignment Research Center research post.If you use WhestBench in academic work, cite the companion paper:Wilson Wu, Victor Lecomte, Michael Winer, George Robinson, Jacob Hilton, and Paul Christiano. "Estimating the expected output of wide random MLPs more efficiently than sampling." arXiv:2605.05179, 2026.The challenge is organized by Alignment Research Center in partnership with AIcrowd.Ready to submit your first estimator?Start with the starter kit, validate locally, then submit through AIcrowd. The first useful milestone is not leaderboard rank—it is one valid end-to-end submission.Make a submissionOpen starter kitRead the Rules

#llm#Overview#Why random
Prize
$150,000
Details
🤖
zindi
ACTIVE|Lvl 4/10

GeoAI Aquaculture Pond Identification Challenge by FAO and ITU

#ai
Prize
TBD
Details
🤖
zindi
ACTIVE|Lvl 4/10

GeoAI Aquaculture Pond Identification Challenge by FAO and ITU

#ai
Prize
TBD
Details
🤖
zindi
ACTIVE|Lvl 5/10

July Study Jam Series: African Credit Scoring Challenge

Prize
TBD
Details
🤖
modelscope
ACTIVE|Lvl 6/10

Aura Pioneers大会启幕,共建 Physical AI 时代开放智能生态

AURA PIONEERS 开放硬件先行者大会由 FrontierX 发起,联合浙江大学启真交叉学科创新创业实验室(ZJU X-Lab)、Delta X 迭代未来孵化器、数有力场等伙伴联合举办,围绕开放硬件、Physical AI 与下一代机器人应用展开,数有力场深度参与策划&生态链接。 当 AI 走出屏幕,第一批开放硬件先行者正在抵达现场过去一个月,6 支入围团队基于 Aura 平台,围绕家庭陪伴、桌面游戏、儿童看护、宠物守护、生活助理等方向展开开发,在真实硬件上完成了一次从想法到 Demo 的验证。8 月 9 日,AURA PIONEERS 开放硬件先行者大会将正式开启。我们邀请你来到现场,一起见证这些开发者如何把 AI 创意带入物理世界。关于Aura Pioneers大会由 FrontierX 发起,联合浙江大学启真交叉学科创新创业实验室(ZJU X-Lab)、Delta X 迭代未来孵化器、数有力场等伙伴联合举办,围绕开放硬件、Physical AI 与下一代机器人应用展开,数有力场深度参与策划&生态链接。本次大会将围绕技术趋势、产业实践、创新项目与生态共创展开,汇聚 大模型、云计算、IoT、高校创新、智能硬件及产业生态伙伴,共同讨论 Physical AI 时代开放创新的新机会。为什么值得来现场?1. 看见 Physical AI 的真实样子Physical AI 不只是一个概念。它正在通过一个个具体 Demo,进入家庭、游戏、陪伴、照护、宠物和生活助理等真实场景。你将看到 AI 如何从文本和屏幕,走向机器、空间和人与人的关系。2. 观察开放硬件如何激发开发者创造一台开放硬件的价值,不只在于它本身能做什么,也在于开发者能基于它创造什么。这次大会将呈现 6 支团队的开发思路、场景判断、Demo 设计与产品想象,帮助你理解开放硬件如何成为 AI 应用创新的新入口。3. 连接开发者、高校与产业生态现场将汇聚开发者团队、高校创新力量、AI 技术平台、硬件与 IoT 伙伴、投资孵化机构及行业观察者。如果你关注 AI 硬件、机器人、Agent、多模态交互、智能终端、开发者生态或早期项目孵化,这里会是一次高密度的交流现场。4. 共同讨论下一代机器人应用下一代机器人应用会从哪里长出来?大会现场将围绕开放硬件、Physical AI、开发者共创与真实场景展开交流,共同探索 AI 进入物理世界后的新机会。适合谁来?欢迎你来到现场:AI / 机器人 / 智能硬件方向开发者;关注 Physical AI、Agent、多模态交互的技术从业者;高校创新团队、学生开发者、科研转化项目成员;智能硬件、IoT、云计算、大模型、语音交互、智能显示等产业伙伴;关注 AI 硬件投资、孵化与商业化的机构代表;正在寻找下一代 AI 应用机会的创业者、产品经理与行业观察者。8 月 9 日,欢迎来到现场。我们相信,未来不是被预测的,而是被体验和连接的期待与未来的超级个体们相遇!Sure AI Field(主办方原名:数有引力·Sure)-数有力场·Sure Ai -起源于一个服务于AI创业者与爱好者的价值平台,连接人文科技。每月组织一场大咖沙龙,周度轻量主题聚会,不仅邀请技术先行者分享前沿洞察,更致力于让硬科技被理解、让创新被看见。数有力场 长期观察AI 如何进入真实生活,持续记录 AI 新物、生活场景与产业变化。如果你是 AI 硬件/具身机器人、生活方式或消费科技品牌,想把产品讲给真实用户听,欢迎联系产品观察、内容或叙事共创。如果你是传统企业、产业机构或创新团队,正在思考如何与 AI 新科技结合、找到更适配自身业务的 AI 生态,也欢迎交流。

#ai
Prize
TBD
Details
🤖
modelscope
ACTIVE|Lvl 6/10

Aura Pioneers大会启幕,共建 Physical AI 时代开放智能生态

AURA PIONEERS 开放硬件先行者大会由 FrontierX 发起,联合浙江大学启真交叉学科创新创业实验室(ZJU X-Lab)、Delta X 迭代未来孵化器、数有力场等伙伴联合举办,围绕开放硬件、Physical AI 与下一代机器人应用展开,数有力场深度参与策划&生态链接。 当 AI 走出屏幕,第一批开放硬件先行者正在抵达现场过去一个月,6 支入围团队基于 Aura 平台,围绕家庭陪伴、桌面游戏、儿童看护、宠物守护、生活助理等方向展开开发,在真实硬件上完成了一次从想法到 Demo 的验证。8 月 9 日,AURA PIONEERS 开放硬件先行者大会将正式开启。我们邀请你来到现场,一起见证这些开发者如何把 AI 创意带入物理世界。关于Aura Pioneers大会由 FrontierX 发起,联合浙江大学启真交叉学科创新创业实验室(ZJU X-Lab)、Delta X 迭代未来孵化器、数有力场等伙伴联合举办,围绕开放硬件、Physical AI 与下一代机器人应用展开,数有力场深度参与策划&生态链接。本次大会将围绕技术趋势、产业实践、创新项目与生态共创展开,汇聚 大模型、云计算、IoT、高校创新、智能硬件及产业生态伙伴,共同讨论 Physical AI 时代开放创新的新机会。为什么值得来现场?1. 看见 Physical AI 的真实样子Physical AI 不只是一个概念。它正在通过一个个具体 Demo,进入家庭、游戏、陪伴、照护、宠物和生活助理等真实场景。你将看到 AI 如何从文本和屏幕,走向机器、空间和人与人的关系。2. 观察开放硬件如何激发开发者创造一台开放硬件的价值,不只在于它本身能做什么,也在于开发者能基于它创造什么。这次大会将呈现 6 支团队的开发思路、场景判断、Demo 设计与产品想象,帮助你理解开放硬件如何成为 AI 应用创新的新入口。3. 连接开发者、高校与产业生态现场将汇聚开发者团队、高校创新力量、AI 技术平台、硬件与 IoT 伙伴、投资孵化机构及行业观察者。如果你关注 AI 硬件、机器人、Agent、多模态交互、智能终端、开发者生态或早期项目孵化,这里会是一次高密度的交流现场。4. 共同讨论下一代机器人应用下一代机器人应用会从哪里长出来?大会现场将围绕开放硬件、Physical AI、开发者共创与真实场景展开交流,共同探索 AI 进入物理世界后的新机会。适合谁来?欢迎你来到现场:AI / 机器人 / 智能硬件方向开发者;关注 Physical AI、Agent、多模态交互的技术从业者;高校创新团队、学生开发者、科研转化项目成员;智能硬件、IoT、云计算、大模型、语音交互、智能显示等产业伙伴;关注 AI 硬件投资、孵化与商业化的机构代表;正在寻找下一代 AI 应用机会的创业者、产品经理与行业观察者。8 月 9 日,欢迎来到现场。我们相信,未来不是被预测的,而是被体验和连接的期待与未来的超级个体们相遇!Sure AI Field(主办方原名:数有引力·Sure)-数有力场·Sure Ai -起源于一个服务于AI创业者与爱好者的价值平台,连接人文科技。每月组织一场大咖沙龙,周度轻量主题聚会,不仅邀请技术先行者分享前沿洞察,更致力于让硬科技被理解、让创新被看见。数有力场 长期观察AI 如何进入真实生活,持续记录 AI 新物、生活场景与产业变化。如果你是 AI 硬件/具身机器人、生活方式或消费科技品牌,想把产品讲给真实用户听,欢迎联系产品观察、内容或叙事共创。如果你是传统企业、产业机构或创新团队,正在思考如何与 AI 新科技结合、找到更适配自身业务的 AI 生态,也欢迎交流。

#ai
Prize
TBD
Details
🤖
civitai
ACTIVE|Lvl 3/10

Meet the Creator Studio: Your Numbers, Finally

One home for every number behind your work - earnings broken out by source, real analytics, and bulk fee & access management. Open to every creator now.

Prize
TBD
Details
🤖
zindi
ACTIVE|Lvl 4/10

Google WAXAL ASR Challenge

Prize
$10 000 USD
Details
🤖
zindi
ACTIVE|Lvl 5/10

Google WAXAL ASR Challenge

Prize
$10 000 USD
Details
🤖
zindi
ACTIVE|Lvl 7/10

University Hackathon by Ngao Labs: Financial Inclusion by Ngao Labs

Prize
TBD
Details
🤖
modelscope
ACTIVE|Lvl 6/10

外滩黑客松·AI Coding 大赛 x Qoder 赛道

Al时代,编程不再是少数人的特权。无论你是学生、上班族、还是退休老人,只要你有想法,就能用AI把它变成真实可用的应用—这就是「外滩黑客松•Al Coding大赛」。这是首个面向全民、不限编程方式的AI编程挑战赛。

#ai
Prize
丰厚奖金池
Details
🤖
modelscope
ACTIVE|Lvl 6/10

AI+♾️ 开发者创作大赛

AI+∞ 开发者创作大赛由魔搭社区与 Qoder 联合发起。每月围绕一个主题,邀请开发者用 AI 把灵感变成可运行的作品,并部署至魔搭创空间,让创意可以被真实体验与分享。第一期聚焦「AI+运营」——运营人,造自己的工具! AI+∞ 开发者创作大赛由魔搭社区与 Qoder 联合发起。每月围绕一个主题,邀请开发者用 AI 把灵感变成可运行的作品,并部署至魔搭创空间,让创意可以被真实体验与分享。第一期聚焦「AI+运营」——运营人,造自己的工具!那些工作中反复出现、耗时又让人抓狂的问题,都可以成为一个作品的起点。动手造吧!也期待未来能和优秀的你继续搭档,从参赛者变成「运营合伙人」,一起共创接下来的比赛。

#应用开发赛#ai
Prize
2万元奖金+Qoder Credit+GPU时长
Details
🤖
modelscope
ACTIVE|Lvl 6/10

Avernet 开源技术沙龙

Agent 已经不只是“会聊天”的工具了,它正在变成能理解目标、调用工具、完成任务的执行伙伴。但在真实场景里,一个 Agent 往往还不够:写一篇热点稿可能需要选题、资料整理、撰稿和审核协作;服务一家小微商家,可能涉及经营分析、活动策划、内容生成和用户触达。问题来了:多个 Agent能不能像一个小团队一样,分工、沟通、协作,把一件事真正做完?这正是 Avernet 想和开发者一起探索的方向。本次沙龙不会只讲概念,我们会结合落地Demo,聊聊多 Agent如何被发现、达成共识、高效执行、自我进化,并围绕同 一个目标协同工作。

Prize
TBD
Details
🤖
civitai
ACTIVE|Lvl 5/10

⚡ Krea 2 Training Contest + 50% Off Training

Krea 2 LoRA training is 50% off, and we're running a Krea 2 Training Contest with 1.8 million Buzz in prizes across four LoRA categories.

#ai
Prize
Enter the Contest
Details
🤖
modelscope
ACTIVE|Lvl 6/10

GOAI 世界人工智能开源大赛

GOAI世界人工智能开源大赛(以下简称“大赛”)由杭州市开源人工智能基金会主办,Agentic AI Foundation(AAIF)和LF AI & Data Foundation 作为全球开源合作伙伴参与支持,并由之江实验室、阿里云智能集团、蚂蚁科技集团股份有限公司、杭州智加未来科技有限公司、云深处科技股份有限公司、杭州群核信息技术有限公司等核心单位承办。大赛面向全球开发者、开源贡献者、高校科研团队、企业 AI 团队、科研机构、创业团队及 AI Builders,聚焦 AI 前沿项目发现、工程验证、开放协作与生态连接。 从杭州出发,面向全球AI Builders,寻找可运行、可验证、可共建的下一代 AI 创新项目!2026年7月16日,GOAI 世界人工智能开源大赛(Global Open-source AI Challenge)官方网站 goaihz.com 正式上线,报名通道面向全球开放。GOAI世界人工智能开源大赛(以下简称“大赛”)由杭州市开源人工智能基金会主办,Agentic AI Foundation(AAIF)和LF AI & Data Foundation 作为全球开源合作伙伴参与支持,并由之江实验室、阿里云智能集团、蚂蚁科技集团股份有限公司、杭州智加未来科技有限公司、云深处科技股份有限公司、杭州群核信息技术有限公司等核心单位承办。大赛面向全球开发者、开源贡献者、高校科研团队、企业 AI 团队、科研机构、创业团队及 AI Builders,聚焦 AI 前沿项目发现、工程验证、开放协作与生态连接。大赛以“Open. Share. Build.” 为核心口号,围绕 新智基座(Agent Infra)、无界应用(Boundless Agents)、前沿探索(AI for Research)、具身未来(Embodied Future)四大赛道展开,重点关注 AI 项目的工程可验证性、开放协作价值、真实场景价值和长期成长潜力,寻找具备可运行 Demo、可复用成果和真实应用潜力的 AI 创新项目。<< 滑动查看更多赛道信息 >>大赛规划总奖金池为500万元级现金激励,设有全场大奖、四大赛道冠亚季军及多个单项奖。其中,全场大奖奖金为 100万元。除现金奖励外,优秀团队还将有机会获得算力与云资源、导师辅导、技术社区曝光、产业场景对接、资本机构直达等多维度资源支持,助力优秀开源 AI 项目持续成长与落地转化。大赛计划于7月21日在杭州云谷中心(杭州市西湖区灯彩街1009号)举行线下启动仪式。届时,组委会将围绕赛事机制、四大赛道、参赛路径、作品提交要求和评审方向进行集中解读,并同步开启作品提交通道。此后,大赛将在北京、上海、深圳、杭州、硅谷、新加坡等城市开展全球城市宣讲,面向四大赛道持续招募优秀团队。按照赛程安排,7月16日大赛官网正式发布并开启报名通道,7月21日举行线下启动仪式,同步开放作品提交通道。7月下旬至8月中旬进入初赛作品提交及全球城市宣讲阶段;8月中旬至下旬开展初赛评审与复赛辅导;8月下旬至9月初进行复赛评审,遴选决赛队伍(各赛道具体安排以官网公布为准)。9月20日起启动决赛周预热,9月22日四大赛道决赛队伍将在杭州展开线下角逐,通过项目路演、Demo展示和专家答辩等环节评选最终奖项。9月23日举办“GOAI DAY”,集中呈现冠军项目Showtime、全场大奖评选及颁奖典礼,展示全球AI开源创新成果。GOAI世界人工智能开源大赛现已面向全球开放报名。欢迎全球开发者、开源贡献者、高校科研团队、企业 AI 团队、科研机构、创业团队及 AI Builders登录官网 goaihz.com 报名参赛,共同构建开放、共享、共建的全球 AI 开源新生态。赛事官网:https://www.goaihz.com/

#应用开发赛#ai
Prize
500万奖金池
Details
🤖
zindi
ACTIVE|Lvl 4/10

Bias Bounty Mapping Equity Challenge

Prize
$10 000 USD
Details
🤖
modelscope
ACTIVE|Lvl 6/10

Production AI Skills 大赛——从「智能体工作流」迈向「生产力 Skills」

第三期(本期,Production AI Skills 大赛):我们再进一步,把目标锁定在「生产力」——鼓励开发者面向生产力级生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等)集成本地 Skill,让端侧智能真正嵌入日常的工作流与生产环节,产出可复用、可商用、有真实价值的技能资产。 一、承前启后:第三期,为「生产力」而来Agent Skills 征文大赛已成功举办两期,我们共同见证了端侧 AI 从概念走向实践的跨越:第一期(3.22–4.30,OpenClaw Skills 挑战赛):聚焦生成式 AI 工具(VLM / 语音识别 / 语音合成 / 图像生成)的本地封装,完成从 0 到 1 的「单点能力封装」。第二期(5.1–6.30,Agent Skills 大赛):升级为面向本地可用的 Skill 封装,并首次要求「基于本地小模型搭建智能体应用」,推动开发者构建完整的 Agentic 工作流。第三期(本期,Production AI Skills 大赛):我们再进一步,把目标锁定在「生产力」——鼓励开发者面向生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等)集成本地 Skill,让端侧智能真正嵌入日常的工作流与生产环节,产出可复用、可商用、有真实价值的技能资产。二、活动背景:Agentic AI 与 Hybrid AI 的端侧生产力新纪元AI 正在经历关键的范式转移:它不再只是对话机器人,而是能够理解意图、自主规划路径、调用外部工具并产生实际后果的「智能体」(Agentic AI)。与此同时,随着端侧算力的爆发,Hybrid AI(混合 AI)架构已成大势——云端处理超大规模逻辑,而将高频响应、隐私敏感与个性化强的任务下沉至 AI PC 处理。英特尔(Intel)作为 AI PC 时代的领航者,通过异构算力(CPU + GPU + NPU)的深度整合,为 Agentic AI 的落地提供坚实的物理基础;魔搭社区(ModelScope)作为开发者生态的摇篮,拥有丰富的模型储备。本期活动,我们希望把这份能力沉淀为「生产力」:将逻辑复杂但体积精炼(≤ 35B)的模型部署在本地,既极大降低推理成本,又确保用户隐私数据「不出机」,并让 Skill 直接服务于真实的生产力场景。三、征文主题与硬性约束1. 核心主题面向生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等)集成一项本地 AI 工具调用(OCR / ASR / TTS / RAG / 数据分析…),解决真实生产力场景的需求,最终产出可被复用、可商用的 Agent Skill。2. 技术约束部署方式:推荐使用Client/Server的模型服务部署方式,实现skill随时调用。运行环境:Skill 中涉及的 AI 模型必须支持纯本地运行(Localhost)。工具适配:Skill 需适配生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等),能够被其作为技能稳定调用。推理框架:推荐使用 OpenVINO™(及其生态工具如 Optimum-intel)构建本地 AI 工具,以充分释放 GPU、NPU 潜力。验证基准:需使用 生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等)作为 Skill 是否可被 Agent 大脑稳定调用的基准测试环境。官方指南:AIPC local skill参考标准: https://github.com/openvino-dev-samples/local-ai-skill-authoring四、推荐方向与场景参考本期尤其鼓励「面向生产力」的 Skill——能嵌入真实工作流、可持续复用、具备商用潜力的方向:推荐方向生产力场景示例Agentic / Hybrid 价值办公提效本地会议纪要自动提取、PPT 一键风格改稿、Excel 复杂公式智能填充。零延迟响应,确保企业内部敏感会议数据绝对安全。开发辅助本地代码 Review、Git Commit 自动生成、API 文档反向生成(可直接在 Qoder 中执行)。离线状态下的高效编程,保护核心算法不外泄。创作创意短视频脚本生成、公众号/小红书文案适配、本地图文自动化排版。端云协同:云端搜集素材,本地用 35B 模型深度二次创作。知识管理个人 PDF/笔记库 RAG、研报摘要提取、本地私人知识库智能问答。构建「永不掉线」且完全私有的个人数字第二大脑。数据分析CSV 自然语言查询、本地数据可视化分析、系统日志异常自动归因。直接读取本地大容量数据集,规避云端流量成本。参考资源:扫码访问英特尔 AI PC 专区,获取更多开发工具、技术文档与实战案例。扫描进入比赛群,实时赛事信息同步、技术交流、问题答疑等。扫码访问英特尔 AI PC 专区,获取更多开发工具、技术文档与实战案例,助力您的 AI 应用开发之旅五、参赛流程:从创意到生产力落地的四步走编写本地 AI 工具:推荐使用 OpenVINO 进行量化与异构加速优化,充分调动 GPU / NPU。生成并验证 Skill:在魔搭 Skills 中心规范下封装技能,并在生产力工具(Qoder、WorkBuddy、TRAE Work 等)的环境下完成指令测试。发布作品包:在魔搭 Skills 中心发布,添加「AI PC」自定义标签;作品需含代码、文档及测试用例。提交技术文章:在魔搭研习社发表文章,沉淀实践路径、在生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等)跑通skill的完整截图/录屏等形式的记录、优化心得与 Hybrid AI 的思考,并添加「Intel AI PC」专题标签。六、评分标准(100%)第三期评分进一步向「生产力价值」倾斜——新增「商用生产力」维度,强调工具的真实使用与场景落地:维度权重说明场景价值30%解决问题的真实性、生产力场景的落地深度、用户群体广度商用生产力30%商用/量产潜力、稳定性与可维护性、能否嵌入真实生产工作流工具使用20%对 Qoder、WorkBuddy、TRAE Work 等生产力 Agent 工具的集成质量、OpenVINO/GPU/NPU 优化、工程实现文章质量10%结构清晰度、可复现性、教学价值创新性10%思路新颖度、与已有方案的差异化传播附加分5%将作品框架截图、流程图或 Skill 成果等图文信息,连同魔搭研习社文章链接、Skill 链接发布至小红书,@OpenVINO中文社区 与 @魔搭ModelScope社区,并加话题 #英特尔 #openvino #魔搭 #agentic #skills 。截止8月31日的阅读量(研习社文章、Skills、小红书累计)超过1000次,得附加分5分。七、丰厚激励实物奖励:前 50 名完整提交作品,即可领取 OpenVINO / 魔搭社区限量周边礼品。现金大奖:TOP 10 作品各获得 1000 元(含税)。生态推广:入选《AI PC Skills Collections》,获得魔搭及英特尔官方全渠道流量扶持。合作池入驻:优秀开发者优先进入「Intel ISV 生态合作伙伴池」,获得后续商业合作与资源对接机会。 让 AI 不止于云端,让智能成为生产力。英特尔与魔搭社区,期待与您共同定义 AI PC 的「生产力灵魂技能」!

#应用开发赛#ai
Prize
10000元奖金+10000元礼品
Details
🤖
aicrowd
ACTIVE|Lvl 6/10

Test P

Test P Predict Wine Quality Predict Wine Quality

##classroom
Prize
TBD
Details
🤖
aicrowd
ACTIVE|Lvl 6/10

Test P

Test P Predict Wine Quality Predict Wine Quality

##classroom
Prize
TBD
Details
🤖
aicrowd
ACTIVE|Lvl 4/10

Trajnet++ (A Trajectory Forecasting Challenge)

Trajnet++ (A Trajectory Forecasting Challenge)

Prize
TBD
Details
🤖
aicrowd
ACTIVE|Lvl 6/10

Trajnet++ (A Trajectory Forecasting Challenge)

Trajnet++ (A Trajectory Forecasting Challenge)

Prize
TBD
Details
🤖
aicrowd
ACTIVE|Lvl 7/10

MEDIQA 2021 - Question Summarization (QS)

MEDIQA 2021 - Question Summarization (QS) ACL-BioNLP Shared Task ACL-BioNLP Shared Task

#nlp#Weight: 1.0
Prize
TBD
Details
🤖
aicrowd
ACTIVE|Lvl 6/10

MEDIQA 2021 - Question Summarization (QS)

MEDIQA 2021 - Question Summarization (QS) ACL-BioNLP Shared Task ACL-BioNLP Shared Task

#nlp#Weight: 1.0
Prize
TBD
Details
🤖
aicrowd
ACTIVE|Lvl 4/10

ECCV 2020 Commands 4 Autonomous Vehicles

ECCV 2020 Commands 4 Autonomous Vehicles

Prize
TBD
Details
🤖
aicrowd
ACTIVE|Lvl 3/10

Multi-Agent Reinforcement Learning for Iterative Reasoning

Multi-Agent Reinforcement Learning for Iterative Reasoning

##reinforcement_learning
Prize
TBD
Details
🤖
aicrowd
ACTIVE|Lvl 3/10

ECCV 2020 Commands 4 Autonomous Vehicles

ECCV 2020 Commands 4 Autonomous Vehicles

Prize
TBD
Details
🤖
aicrowd
ACTIVE|Lvl 6/10

Multi-Agent Reinforcement Learning for Iterative Reasoning

Multi-Agent Reinforcement Learning for Iterative Reasoning

##reinforcement_learning
Prize
TBD
Details
🤖
modelscope
ACTIVE|Lvl 3/10

AI for Science 闭门沙龙

《浦江科技评论》·最创社区 × 上海天使会 x 创新工场 联合主办,魔搭社区合作支持,聚焦 AI4S 范式变革下的中国机会。 《浦江科技评论》·最创社区 × 上海天使会 x 创新工场 联合主办,魔搭社区 合作支持,聚焦 AI4S 范式变革下的中国机会。AI 正在重写科研本身。从材料发现到药物研发,从蛋白质设计到自动化实验室,AI 开始真正参与科学发现的每一个环节。而在这场全球性的科研范式变革中,中国的结构性机会在哪里?活动聚焦— 全球趋势:AI4S 全球发展图景与中国机会— 科研范式:Agentic AI 驱动科学研究— 行业实践:AI4S 在生命科学、材料科学等领域的真实应用案例— 未来展望:AI4S 下一阶段的关键挑战与突破方向嘉宾阵容任博冰 · 创新工场执行董事&前沿科技基金总经理唐诗翔 · 香港中文大学博士后、上海人工智能实验室青年科学家张书铭 · 深度原理创始团队&战略与生态合作负责人王宇光 · 途深智合创始人、上海交通大学副教授郑 文 · 鸿之微科技(上海)股份有限公司战略规划负责人丁 峰 · 创新工场副总裁……特别环节AI4S 创新企业开放麦,欢迎有意分享的企业同步报名📅 8月4日(周二)14:00–18:00📍 上海 · 邀请制沙龙(同步线上分享)欢迎AI+生命科学、脑机接口、材料方向的创业者、研究者及企业研发负责人扫码报名。

#ai
Prize
TBD
Details
🤖
civitai
ACTIVE|Lvl 4/10

Open Your Own Creator Shop

Design and sell custom cosmetics, list your models, and run a real storefront right on your profile. Creator Shops are here - Creator Program members now, everyone soon.

Prize
TBD
Details
🤖
zindi
ACTIVE|Lvl 7/10

July Study Jam Series: African Credit Scoring Challenge

Prize
TBD
Details
🤖
civitai
ACTIVE|Lvl 7/10

Meet the Creator Studio: Your Numbers, Finally

One home for every number behind your work - earnings broken out by source, real analytics, and bulk fee & access management. Open to every creator now.

Prize
TBD
Details
🤖
zindi
ACTIVE|Lvl 7/10

University Hackathon by Ngao Labs: Financial Inclusion by Ngao Labs

Prize
TBD
Details
🤖
modelscope
ACTIVE|Lvl 5/10

外滩黑客松·AI Coding 大赛 x Qoder 赛道

Al时代,编程不再是少数人的特权。无论你是学生、上班族、还是退休老人,只要你有想法,就能用AI把它变成真实可用的应用—这就是「外滩黑客松•Al Coding大赛」。这是首个面向全民、不限编程方式的AI编程挑战赛。

#ai
Prize
丰厚奖金池
Details
🤖
modelscope
ACTIVE|Lvl 7/10

AI+♾️ 开发者创作大赛

AI+∞ 开发者创作大赛由魔搭社区与 Qoder 联合发起。每月围绕一个主题,邀请开发者用 AI 把灵感变成可运行的作品,并部署至魔搭创空间,让创意可以被真实体验与分享。第一期聚焦「AI+运营」——运营人,造自己的工具! AI+∞ 开发者创作大赛由魔搭社区与 Qoder 联合发起。每月围绕一个主题,邀请开发者用 AI 把灵感变成可运行的作品,并部署至魔搭创空间,让创意可以被真实体验与分享。第一期聚焦「AI+运营」——运营人,造自己的工具!那些工作中反复出现、耗时又让人抓狂的问题,都可以成为一个作品的起点。动手造吧!也期待未来能和优秀的你继续搭档,从参赛者变成「运营合伙人」,一起共创接下来的比赛。

#应用开发赛#ai
Prize
2万元奖金+Qoder Credit+GPU时长
Details
🤖
modelscope
ACTIVE|Lvl 7/10

Avernet 开源技术沙龙

Agent 已经不只是“会聊天”的工具了,它正在变成能理解目标、调用工具、完成任务的执行伙伴。但在真实场景里,一个 Agent 往往还不够:写一篇热点稿可能需要选题、资料整理、撰稿和审核协作;服务一家小微商家,可能涉及经营分析、活动策划、内容生成和用户触达。问题来了:多个 Agent能不能像一个小团队一样,分工、沟通、协作,把一件事真正做完?这正是 Avernet 想和开发者一起探索的方向。本次沙龙不会只讲概念,我们会结合落地Demo,聊聊多 Agent如何被发现、达成共识、高效执行、自我进化,并围绕同 一个目标协同工作。

Prize
TBD
Details
🤖
civitai
ACTIVE|Lvl 7/10

⚡ Krea 2 Training Contest + 50% Off Training

Krea 2 LoRA training is 50% off, and we're running a Krea 2 Training Contest with 1.8 million Buzz in prizes across four LoRA categories.

#ai
Prize
Enter the Contest
Details
🤖
modelscope
ACTIVE|Lvl 6/10

GOAI 世界人工智能开源大赛

GOAI世界人工智能开源大赛(以下简称“大赛”)由杭州市开源人工智能基金会主办,Agentic AI Foundation(AAIF)和LF AI & Data Foundation 作为全球开源合作伙伴参与支持,并由之江实验室、阿里云智能集团、蚂蚁科技集团股份有限公司、杭州智加未来科技有限公司、云深处科技股份有限公司、杭州群核信息技术有限公司等核心单位承办。大赛面向全球开发者、开源贡献者、高校科研团队、企业 AI 团队、科研机构、创业团队及 AI Builders,聚焦 AI 前沿项目发现、工程验证、开放协作与生态连接。 从杭州出发,面向全球AI Builders,寻找可运行、可验证、可共建的下一代 AI 创新项目!2026年7月16日,GOAI 世界人工智能开源大赛(Global Open-source AI Challenge)官方网站 goaihz.com 正式上线,报名通道面向全球开放。GOAI世界人工智能开源大赛(以下简称“大赛”)由杭州市开源人工智能基金会主办,Agentic AI Foundation(AAIF)和LF AI & Data Foundation 作为全球开源合作伙伴参与支持,并由之江实验室、阿里云智能集团、蚂蚁科技集团股份有限公司、杭州智加未来科技有限公司、云深处科技股份有限公司、杭州群核信息技术有限公司等核心单位承办。大赛面向全球开发者、开源贡献者、高校科研团队、企业 AI 团队、科研机构、创业团队及 AI Builders,聚焦 AI 前沿项目发现、工程验证、开放协作与生态连接。大赛以“Open. Share. Build.” 为核心口号,围绕 新智基座(Agent Infra)、无界应用(Boundless Agents)、前沿探索(AI for Research)、具身未来(Embodied Future)四大赛道展开,重点关注 AI 项目的工程可验证性、开放协作价值、真实场景价值和长期成长潜力,寻找具备可运行 Demo、可复用成果和真实应用潜力的 AI 创新项目。<< 滑动查看更多赛道信息 >>大赛规划总奖金池为500万元级现金激励,设有全场大奖、四大赛道冠亚季军及多个单项奖。其中,全场大奖奖金为 100万元。除现金奖励外,优秀团队还将有机会获得算力与云资源、导师辅导、技术社区曝光、产业场景对接、资本机构直达等多维度资源支持,助力优秀开源 AI 项目持续成长与落地转化。大赛计划于7月21日在杭州云谷中心(杭州市西湖区灯彩街1009号)举行线下启动仪式。届时,组委会将围绕赛事机制、四大赛道、参赛路径、作品提交要求和评审方向进行集中解读,并同步开启作品提交通道。此后,大赛将在北京、上海、深圳、杭州、硅谷、新加坡等城市开展全球城市宣讲,面向四大赛道持续招募优秀团队。按照赛程安排,7月16日大赛官网正式发布并开启报名通道,7月21日举行线下启动仪式,同步开放作品提交通道。7月下旬至8月中旬进入初赛作品提交及全球城市宣讲阶段;8月中旬至下旬开展初赛评审与复赛辅导;8月下旬至9月初进行复赛评审,遴选决赛队伍(各赛道具体安排以官网公布为准)。9月20日起启动决赛周预热,9月22日四大赛道决赛队伍将在杭州展开线下角逐,通过项目路演、Demo展示和专家答辩等环节评选最终奖项。9月23日举办“GOAI DAY”,集中呈现冠军项目Showtime、全场大奖评选及颁奖典礼,展示全球AI开源创新成果。GOAI世界人工智能开源大赛现已面向全球开放报名。欢迎全球开发者、开源贡献者、高校科研团队、企业 AI 团队、科研机构、创业团队及 AI Builders登录官网 goaihz.com 报名参赛,共同构建开放、共享、共建的全球 AI 开源新生态。赛事官网:https://www.goaihz.com/

#应用开发赛#ai
Prize
500万奖金池
Details
🤖
zindi
ACTIVE|Lvl 6/10

Bias Bounty Mapping Equity Challenge

Prize
$10 000 USD
Details
🤖
modelscope
ACTIVE|Lvl 4/10

Production AI Skills 大赛——从「智能体工作流」迈向「生产力 Skills」

第三期(本期,Production AI Skills 大赛):我们再进一步,把目标锁定在「生产力」——鼓励开发者面向生产力级生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等)集成本地 Skill,让端侧智能真正嵌入日常的工作流与生产环节,产出可复用、可商用、有真实价值的技能资产。 一、承前启后:第三期,为「生产力」而来Agent Skills 征文大赛已成功举办两期,我们共同见证了端侧 AI 从概念走向实践的跨越:第一期(3.22–4.30,OpenClaw Skills 挑战赛):聚焦生成式 AI 工具(VLM / 语音识别 / 语音合成 / 图像生成)的本地封装,完成从 0 到 1 的「单点能力封装」。第二期(5.1–6.30,Agent Skills 大赛):升级为面向本地可用的 Skill 封装,并首次要求「基于本地小模型搭建智能体应用」,推动开发者构建完整的 Agentic 工作流。第三期(本期,Production AI Skills 大赛):我们再进一步,把目标锁定在「生产力」——鼓励开发者面向生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等)集成本地 Skill,让端侧智能真正嵌入日常的工作流与生产环节,产出可复用、可商用、有真实价值的技能资产。二、活动背景:Agentic AI 与 Hybrid AI 的端侧生产力新纪元AI 正在经历关键的范式转移:它不再只是对话机器人,而是能够理解意图、自主规划路径、调用外部工具并产生实际后果的「智能体」(Agentic AI)。与此同时,随着端侧算力的爆发,Hybrid AI(混合 AI)架构已成大势——云端处理超大规模逻辑,而将高频响应、隐私敏感与个性化强的任务下沉至 AI PC 处理。英特尔(Intel)作为 AI PC 时代的领航者,通过异构算力(CPU + GPU + NPU)的深度整合,为 Agentic AI 的落地提供坚实的物理基础;魔搭社区(ModelScope)作为开发者生态的摇篮,拥有丰富的模型储备。本期活动,我们希望把这份能力沉淀为「生产力」:将逻辑复杂但体积精炼(≤ 35B)的模型部署在本地,既极大降低推理成本,又确保用户隐私数据「不出机」,并让 Skill 直接服务于真实的生产力场景。三、征文主题与硬性约束1. 核心主题面向生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等)集成一项本地 AI 工具调用(OCR / ASR / TTS / RAG / 数据分析…),解决真实生产力场景的需求,最终产出可被复用、可商用的 Agent Skill。2. 技术约束部署方式:推荐使用Client/Server的模型服务部署方式,实现skill随时调用。运行环境:Skill 中涉及的 AI 模型必须支持纯本地运行(Localhost)。工具适配:Skill 需适配生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等),能够被其作为技能稳定调用。推理框架:推荐使用 OpenVINO™(及其生态工具如 Optimum-intel)构建本地 AI 工具,以充分释放 GPU、NPU 潜力。验证基准:需使用 生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等)作为 Skill 是否可被 Agent 大脑稳定调用的基准测试环境。官方指南:AIPC local skill参考标准: https://github.com/openvino-dev-samples/local-ai-skill-authoring四、推荐方向与场景参考本期尤其鼓励「面向生产力」的 Skill——能嵌入真实工作流、可持续复用、具备商用潜力的方向:推荐方向生产力场景示例Agentic / Hybrid 价值办公提效本地会议纪要自动提取、PPT 一键风格改稿、Excel 复杂公式智能填充。零延迟响应,确保企业内部敏感会议数据绝对安全。开发辅助本地代码 Review、Git Commit 自动生成、API 文档反向生成(可直接在 Qoder 中执行)。离线状态下的高效编程,保护核心算法不外泄。创作创意短视频脚本生成、公众号/小红书文案适配、本地图文自动化排版。端云协同:云端搜集素材,本地用 35B 模型深度二次创作。知识管理个人 PDF/笔记库 RAG、研报摘要提取、本地私人知识库智能问答。构建「永不掉线」且完全私有的个人数字第二大脑。数据分析CSV 自然语言查询、本地数据可视化分析、系统日志异常自动归因。直接读取本地大容量数据集,规避云端流量成本。参考资源:扫码访问英特尔 AI PC 专区,获取更多开发工具、技术文档与实战案例。扫描进入比赛群,实时赛事信息同步、技术交流、问题答疑等。扫码访问英特尔 AI PC 专区,获取更多开发工具、技术文档与实战案例,助力您的 AI 应用开发之旅五、参赛流程:从创意到生产力落地的四步走编写本地 AI 工具:推荐使用 OpenVINO 进行量化与异构加速优化,充分调动 GPU / NPU。生成并验证 Skill:在魔搭 Skills 中心规范下封装技能,并在生产力工具(Qoder、WorkBuddy、TRAE Work 等)的环境下完成指令测试。发布作品包:在魔搭 Skills 中心发布,添加「AI PC」自定义标签;作品需含代码、文档及测试用例。提交技术文章:在魔搭研习社发表文章,沉淀实践路径、在生产力级 AI Agent 工具(Qoder、WorkBuddy、TRAE Work 等)跑通skill的完整截图/录屏等形式的记录、优化心得与 Hybrid AI 的思考,并添加「Intel AI PC」专题标签。六、评分标准(100%)第三期评分进一步向「生产力价值」倾斜——新增「商用生产力」维度,强调工具的真实使用与场景落地:维度权重说明场景价值30%解决问题的真实性、生产力场景的落地深度、用户群体广度商用生产力30%商用/量产潜力、稳定性与可维护性、能否嵌入真实生产工作流工具使用20%对 Qoder、WorkBuddy、TRAE Work 等生产力 Agent 工具的集成质量、OpenVINO/GPU/NPU 优化、工程实现文章质量10%结构清晰度、可复现性、教学价值创新性10%思路新颖度、与已有方案的差异化传播附加分5%将作品框架截图、流程图或 Skill 成果等图文信息,连同魔搭研习社文章链接、Skill 链接发布至小红书,@OpenVINO中文社区 与 @魔搭ModelScope社区,并加话题 #英特尔 #openvino #魔搭 #agentic #skills 。截止8月31日的阅读量(研习社文章、Skills、小红书累计)超过1000次,得附加分5分。七、丰厚激励实物奖励:前 50 名完整提交作品,即可领取 OpenVINO / 魔搭社区限量周边礼品。现金大奖:TOP 10 作品各获得 1000 元(含税)。生态推广:入选《AI PC Skills Collections》,获得魔搭及英特尔官方全渠道流量扶持。合作池入驻:优秀开发者优先进入「Intel ISV 生态合作伙伴池」,获得后续商业合作与资源对接机会。 让 AI 不止于云端,让智能成为生产力。英特尔与魔搭社区,期待与您共同定义 AI PC 的「生产力灵魂技能」!

#应用开发赛#ai
Prize
10000元奖金+10000元礼品
Details
🤖
drivendata
ACTIVE|Lvl 4/10

Pump it Up: Data Mining the Water Table

development

Prize
TBD
Details
🤖
drivendata
ACTIVE|Lvl 5/10

DengAI: Predicting Disease Spread

health

#ai
Prize
TBD
Details
🤖
drivendata
ACTIVE|Lvl 4/10

Richter's Predictor: Modeling Earthquake Damage

disasters

Prize
TBD
Details
🤖
drivendata
ACTIVE|Lvl 5/10

Conser-vision Practice Area: Image Classification

climate

Prize
TBD
Details
...