LearnPrompt vs Surge AI

Side-by-side comparison of features, pricing, and ratings

Analysis reviewed Live tool data as of 2026-09-29
Cross-checked through our multi-step verification ·
Saved

At a glance

DimensionLearnPromptSurge AI
What it isFree open-source Chinese AI practice wiki (8 paths, 47 tutorials)Human RLHF data, red teaming and eval benchmarks vendor
PriceFree, permanentlyContact sales; scoping call required
BuyerDevelopers already using Claude Code/CodexFrontier labs and post-training teams
Core deliverableTask cards, CLAUDE.md/AGENTS.md/SKILL.md writing guidesExpert-labeled preference data, adversarial testing, citable benchmarks
Named assetsHarness 五组件排查法; Loop Engineering path; Hermes/OpenClaw guideGDP.pdf, ComplexConstraints, HANDBOOK.md, Chartography, Tuesday Work Index
FormatText tutorials, no video, no certificatesCustom data programs and published benchmark suites

These two don't sit on the same shortlist, and pretending otherwise wastes your time. If you're an individual developer or small team trying to turn scattered Claude Code/Codex usage into a repeatable workflow — task cards, project rules, SKILL.md, loop engineering — LearnPrompt costs nothing and is the only one of the two that speaks to you. If you run post-training at a frontier lab and need credentialed experts producing RLHF preference data, red-team findings, or a benchmark number you can cite in a system card (OpenAI did exactly that with GDP.pdf in GPT-5.6), Surge AI is the vendor and LearnPrompt is irrelevant. Pick based on which sentence described your job.

LearnPrompt
LearnPrompt

永久免费的中文 AI 实战 Wiki,教你把「想法→任务卡→Agent 交付→复盘」跑通真实项目

Visit Website
Surge AI
Surge AI

Expert human RLHF data, red teaming, and citable AI benchmarks for frontier model labs

Visit Website
Pricing
Free
Contact Sales
Plans
$0
—
Popularity
1 views
7.4k views
Skill Level
Intermediate
Advanced
API Available
Platforms
Web
WebAPI
Categories
💻 Code & Development🔬 Research & Education
🏷️ Data Labeling & Training Data
Features
8 条学习路径:AI 编程入门、Claude Code、Codex、Agent 工程等
47 篇教程,每篇标注一手来源与核验日期
任务卡模板:拆出输入、输出、失败态与验收标准
CLAUDE.md / AGENTS.md 项目规则写作指南
SKILL.md 写作指南:把重复流程包装为可复用 Skill
Agentic Coding 最小工作流:计划、改动、验证、复盘
Harness 五组件排查法,定位 Agent 跑偏缺失的那一层
Claude Code 与 Codex 执行面选型对比(CLI / IDE / 桌面 / Cloud)
Loop Engineering:为下一轮留下状态、验收结果和下一步
Obsidian AI 知识工作台:目录、索引与项目交接
Hermes / OpenClaw 长驻 Agent 架构导读与成本核算
可重放 Showcase:正常路径、失败场景与越界反例都要跑
开源课程,社区可贡献,支持 GitHub 协作
在线免费访问,无需注册
旧版 LearnPrompt 站点归档可回溯
Expert human workforce spanning doctors, lawyers, engineers, and writers
RLHF preference data collection and human feedback for model fine-tuning
Red teaming and adversarial testing staffed with credentialled domain specialists
Off-the-shelf post-training runs built on expert evaluation data
SWE consultant network for technical and software engineering tasks
Agentic coding task sets for post-training (1,700 tasks lifted Kimi K2.7 +20.0pp on SWE-Marathon)
GDP.pdf benchmark for real-world professional document comprehension
ComplexConstraints benchmark for entangled, conditional instruction following
HANDBOOK.md benchmark for long-context policy adherence against expert handbooks
Chartography benchmark for professional chart reading: Kaplan-Meier curves, candlesticks, Bode plots
Tuesday Work Index composite benchmark for real professional work capabilities
DAYJOB vertical benchmark suites for economically valuable agents in Healthcare and Finance
Riemann-bench for extreme math verification
EnterpriseBench and CoreCraft RL environments
MCP-native RL environments for enterprise agent tasks
Integrations
Claude Code
Codex
Obsidian
ChatGPT
Hermes
OpenClaw
GitHub

Feature-by-feature

LearnPrompt teaches a workflow; Surge AI produces data. That's the whole gap. LearnPrompt's units are task cards: each one specifies input, output, failure modes and acceptance criteria, so you can hand a job to an Agent, read the diff, run the tests, and decide whether the lesson belongs in AGENTS.md or SKILL.md. Its 8 paths cover AI coding basics, Claude Code, Codex, Agent engineering, Agent Skills, Loop Engineering, Obsidian AI, and Hermes/OpenClaw long-running agent architecture. The genuinely differentiating material is the boring-sounding operational stuff: the Harness 五组件排查法 for locating the missing layer when an Agent drifts, a CLI/IDE/desktop/Cloud selection comparison between Claude Code and Codex, and a rule that Showcases must be replayed including the out-of-bounds counterexamples. Every tutorial lists a first-party source and a verification date. Surge AI is a different species: a credentialed human workforce (doctors, lawyers, engineers, writers) doing RLHF preference collection, red teaming, multimodal and reasoning-heavy labeling, and off-the-shelf post-training runs. Its evaluation catalog is the loudest part — GDP.pdf (professional document comprehension; OpenAI's GPT-5.6 flagship scored 30.7%), ComplexConstraints (entangled, context-inferred instruction following), HANDBOOK.md (long-context policy following), Chartography (Kaplan-Meier curves, Bode plots), the Tuesday Work Index composite, DAYJOB suites for Healthcare/Finance, and Riemann-bench. One is a knowledge base you read; the other is a service that returns labeled data and numbers.

Pricing compared

LearnPrompt is free and permanently so — no tiers, no gating, no upsell path in the data. It's open source on GitHub, maintained by Carl, with old course versions kept archived. Your real cost is time: 47 tutorials, no video, no certificate at the end, and content that assumes you've already run a terminal and Git once. Surge AI can't be priced from the page: pricing_type is contact, and the stated disqualifier for buyers is showing up to a scoping call without a budget and a scoped pilot. That tells you the engagement model — custom programs sized to your labeling volume, workforce specialization, and benchmark needs, negotiated rather than listed. The two numbers never meet. A hobbyist comparing them on cost learns nothing useful; the honest comparison is zero dollars against an enterprise contract that requires procurement. If a listed price is a hard requirement for you, Surge AI is out of scope by definition, and if you're a solo developer, so is its entire buyer profile. Note the one price-adjacent fact in the news: training a 4B model on 1,000 expert-written ComplexConstraints rubrics moved MultiChallenge +10.1 and AdvancedIF +8.4 — efficiency framing aimed at teams weighing data spend against compute.

Who should pick which

  • Developer with scattered Claude Code/Codex habits
    Pick: LearnPrompt

    The task-card format and CLAUDE.md/AGENTS.md/SKILL.md guides exist precisely to convert ad-hoc prompting into a repeatable plan-change-verify-retro loop.

  • Technical lead comparing Claude Code vs Codex execution surfaces
    Pick: LearnPrompt

    It publishes a CLI/IDE/desktop/Cloud selection comparison plus a Harness five-component troubleshooting method — free, and no sales call required.

  • Obsidian-based note taker who wants AI to read their context
    Pick: LearnPrompt

    There is a dedicated Obsidian AI path covering directory structure, indexes and project handoff, which is otherwise thin territory.

  • Post-training lead at a frontier lab needing citable eval numbers
    Pick: Surge AI

    GDP.pdf, ComplexConstraints and the Tuesday Work Index are already cited in other labs' release materials — that is the citation surface you cannot build in-house quickly.

  • AI safety team running red teaming with credentialed specialists
    Pick: Surge AI

    Adversarial testing staffed by domain specialists is a listed core service; a wiki, however good, cannot supply the graders.

Frequently Asked Questions

LearnPrompt vs Surge AI: which should you choose?

These two don't sit on the same shortlist, and pretending otherwise wastes your time. If you're an individual developer or small team trying to turn scattered Claude Code/Codex usage into a repeatable workflow — task cards, project rules, SKILL.md, loop engineering — LearnPrompt costs nothing and is the only one of the two that speaks to you. If you run post-training at a frontier lab and need credentialed experts producing RLHF preference data, red-team findings, or a benchmark number you can cite in a system card (OpenAI did exactly that with GDP.pdf in GPT-5.6), Surge AI is the vendor and LearnPrompt is irrelevant. Pick based on which sentence described your job.

Is LearnPrompt a competitor to Surge AI?

No. One publishes free tutorials about agentic coding workflows; the other sells expert-labeled training data and evaluation benchmarks. Different buyers, different budgets, no overlap.

Can LearnPrompt's material substitute for Surge AI's benchmarks?

No. LearnPrompt teaches you to validate your own agent runs with diffs, tests and replayable showcases. Surge's benchmarks measure frontier model capability against expert-authored rubrics and are cited in release notes and system cards.

Who should not use LearnPrompt?

Complete beginners who haven't run a terminal or Git once, people wanting tool rankings or model leaderboards, video-first learners, and anyone seeking certificates or formal certification.

What disqualifies a Surge AI buyer?

Simple classification or bulk low-complexity labeling, fully automated evaluation with no human graders, needing only a rough internal number, or arriving without a scoped pilot and budget.

Does LearnPrompt cover model selection or rankings?

No — it explicitly excludes tool rankings and marketing-style introductions. The closest it gets is a selection comparison between Claude Code and Codex execution surfaces.

How current is Surge AI's benchmark catalog?

Its news cadence is recent and specific: ComplexConstraints was introduced in August 2026, the Tuesday Work Index followed, and OpenAI cited GDP.pdf in its July 2026 GPT-5.6 release.

More LearnPrompt or Surge AI comparisons

Explore each tool further

Browse these categories

Still deciding? Get the weekly AI tools brief

One email a week — new tools, honest comparisons, no spam.

Last reviewed: September 28, 2026