RadarGraph World Algorithm Safe Build Plan
Planning only. Do not implement this plan, modify product code, launch new crons, restart gateways, push commits, or let worker agents write external state until Hareesh approves the execution scope.
Goal: In 12 hours, produce a founder-grade, technically implementable “world algorithm” for Padawan RadarGraph: a source-intelligence ranking and synthesis system with PageRank-level ambition, using Grok radar, Codex adversarial/protocol engineering, and a compounding evidence graph.
Architecture: Treat this as a research-and-systems design sprint, not a hype sprint. The output is a versioned algorithm spec plus a minimal benchmark harness plan. The system’s core is a graph where sources, claims, critiques, searches, user judgments, and successful tactics become scored nodes/edges that improve future research trajectories.
Tech Stack: Hermes cron/session orchestration, Grok CLI for live radar, Codex CLI for adversarial/system critique, local HTML/Markdown artifacts, optional local SQLite/JSONL prototype spec, web/arXiv/source fetches, no production writes.
Useful anchors: PageRank paper, Codex CLI, Hermes Agent docs, arXiv.
Safety posture
Non-negotiable boundaries
- No production code changes unless Hareesh separately approves the exact repo/path.
- No cross-profile reads/writes beyond this profile’s permitted working surface.
- No credentials, tokens, auth files, or gateway internals.
- No deploys, pushes, PRs, external messages, paid API changes, billing, or account operations.
- No autonomous self-scheduling beyond the approved 12h sprint. If a cron already exists, inspect/pause/update deliberately; do not multiply jobs.
- Every claim about “latest advancement” must be either source-backed or labeled hypothesis.
- Grok output is radar, not truth. Verify load-bearing claims against primary/serious sources where possible.
- Codex output is critique/design input, not authority. Padawan must adjudicate.
- Deliverable is a design + benchmark + prototype plan, not a fake implementation. If something is not built/tested, say so.
Risk model
This sprint can fail in five main ways:
- Grandiosity drift: “beat every product” becomes abstract hype instead of measurable axes.
- Tool overreach: crons/workers multiply, become noisy, or touch unsafe surfaces.
- Source laundering: Grok or web snippets get promoted into facts without receipts.
- Algorithm theater: formulas look serious but do not improve outputs.
- Taste overfit: system learns Hareesh’s prose/taste but not durable research judgment.
The plan below is designed to force concrete outputs, measurable deltas, and stop gates.
Current context / assumptions
Known context
- Hareesh’s intuition: Padawan’s novel lane is not “better AI UI.” It is Grok radar + high-quality synthesis + adversarial compounding.
- The desired impact analogy is PageRank: a deceptively clean ranking principle that changed the quality frontier of an information product.
- Existing draft concept: RadarGraph — live frontier radar, source/claim graph, Codex adversary/protocol engineer, scoring loop.
- A benchmark baseline artifact already exists in the conversation:
grok-radar-pingpong-benchmark-20260811-2213.html. - A first RadarGraph algorithm sketch exists:
padawan-radargraph-algorithm-20260811-2217.html. - A “World Algorithm — RadarGraph 12h build” cron appears scheduled from an interrupted prior turn. Safe handling requires inspecting/updating/pausing rather than blindly adding more automation.
Assumptions
- The target is algorithm/spec quality, not immediate production implementation.
- The safe 12h output is a sequence of versioned artifacts and a final master spec.
- “Every latest advancement” means bounded coverage across the relevant frontier: graph ranking, retrieval/reranking, knowledge graphs, source credibility, agent trajectories, multi-agent critique, uncertainty calibration, active learning/bandits, preference learning, and provenance.
- We should bias toward a system that can be implemented locally first with JSONL/SQLite before any complex infra.
The central product thesis to test
RadarGraph should be to AI research agents what PageRank was to search: not merely a better model call, but a ranking/compounding structure that makes better information surface because the system represents relationships better than competitors do.
The candidate PageRank-like insight:
A research result is valuable when high-credibility sources, independent claims, adversarial critiques, and downstream user decisions mutually reinforce each other over time — and when the graph penalizes ungrounded fluency, source monoculture, unresolved critique, and non-actionable synthesis.
So the “world algorithm” must rank not only sources, but also claims, critique tickets, search trajectories, synthesis versions, user judgments, reusable tactics, and future radar priorities.
Proposed 12-hour sprint architecture
Sprint outputs
By the end of 12 hours, produce:
- Master algorithm spec as HTML + Markdown.
- Data model for the graph.
- Core scoring formulas with rationale and failure modes.
- Run protocol for Grok → Padawan → Codex → targeted radar → synthesis.
- Benchmark suite spec comparing against generic deep research products.
- Prototype build plan for a local harness.
- Safety/quality gates for future automation.
- Open questions / kill criteria that could prove the algorithm is not special.
Workstream roles
- Padawan/controller: scope, synthesis, taste calibration, final adjudication, artifact writing.
- Grok/radar: live frontier scan, source discovery, weird leads, social/product discourse, competitor movement.
- Codex/adversary: protocol critique, algorithmic consistency, data model audit, benchmark loopholes, implementation-shape review.
- Web/source verification: primary source confirmation for load-bearing claims.
- Benchmark harness: define how to prove lift, not just produce a nice document.
Safe execution timeline
Hour 0: Freeze and align
Objective: Prevent unsafe automation and clarify the sprint contract.
Actions:
- Inspect active crons.
- If a World Algorithm cron is active from the interrupted turn, either pause it until Hareesh approves this plan or update it to follow this plan exactly.
- Confirm no new cron should recursively create/modify crons.
- Set all sprint jobs to deliver artifacts only, not chat walls.
- Establish one scratch directory:
/Users/openclaw/.hermes/profiles/personal/tmp/radargraph-world-algo-YYYYMMDD/. - Establish one artifact directory:
/Users/openclaw/.hermes/profiles/personal/brief-viewer/.
Deliverable: short Telegram checkpoint or artifact if long.
Stop gate: If automation state is ambiguous, pause the sprint jobs and proceed manually.
Hour 1: Frontier map
Objective: Build a source-backed map of relevant “latest advancements.”
Source lanes: graph ranking / PageRank successors; neural retrieval, reranking, hybrid search; knowledge graphs and claim graphs; agent trajectory evaluation; multi-agent debate, critique, verifier loops; source credibility and provenance; active learning / bandits for search policy; preference learning / taste calibration; deep research agent benchmarks and competitors.
Actions: Run Grok radar for frontier concepts and named systems; run web/source search for primary anchors; record each lead with source id, title, url, source class, claim relevance, reliability note, and accessed date.
Deliverable: frontier-map-v1.md or HTML section.
Stop gate: No source-backed algorithm claims may enter the master spec yet; this is only the map.
Hour 2: Competitor baseline and attack surface
Objective: Define what “beat them” means in concrete product terms.
Baselines: Generic ChatGPT/Claude/Gemini deep research; Perplexity-style answer engine; Manus/Genspark-style autonomous report agent; Notion/Granola/Limitless-style memory/workflow products; internal Padawan single-pass output.
Quality axes: novel source yield; claim grounding and traceability; adversarial survival; decision usefulness; taste/voice fit without mimicry; trajectory lift between passes; graph reuse in future runs; time/cost/noise efficiency.
Deliverable: baseline matrix.
Stop gate: If “beat them” cannot be measured on an axis, that axis is not allowed in the product claim.
Hour 3: Data model v1
Objective: Design the graph state that compounds.
Core nodes: SourceNode, AuthorNode, ClaimNode, QuestionNode, SearchTrajectoryNode, CritiqueNode, SynthesisNode, UserJudgmentNode, TacticNode, BenchmarkFixtureNode.
Core edges: SUPPORTS, CONTRADICTS, CITES, DERIVED_FROM, CRITIQUES, RESOLVES, FAILED_BY, PREFERRED_OVER, REUSED_IN, UPGRADED_BY.
Deliverable: schema spec with JSON examples.
Stop gate: If the model cannot represent why a future answer got better, it is not a compounding graph.
Hour 4: Ranking formula candidates
Objective: Draft PageRank-like scoring primitives.
Candidate scores:
- SourceRank: credibility × independence × recency × proximity to primary evidence × historical usefulness.
- ClaimRank: support strength × source independence × contradiction penalty × critique survival × decision leverage.
- CritiqueRank: severity × target load-bearingness × specificity × falsifiability × historical yield.
- TrajectoryRank: novelty yield × claim flips × critique resolution × time cost × downstream reuse.
- TasteFitScore: Hareesh preference signal × blind win rate × anti-slop penalty × non-mimicry constraint.
- ActionRank: expected founder decision impact × reversibility × cost × speed × evidence confidence.
Deliverable: formulas plus examples.
Stop gate: Every formula must name a gaming/failure mode.
Hour 5: Codex adversarial protocol audit
Objective: Treat the algorithm like software and attack it.
Codex prompt should ask: Where can this scoring system be gamed? Which variables are unobservable or fake precision? Which graph updates create feedback loops or taste overfit? What minimal tests would catch algorithm theater? What data model will break first? Which ranking formula has the highest risk of Goodharting?
Deliverable: Codex critique receipt + accepted/rejected findings.
Stop gate: Do not accept Codex critique wholesale. Padawan must classify each finding as accepted, partially accepted, rejected, or needs evidence.
Hour 6: RadarGraph v2 synthesis
Objective: Merge frontier map, baseline matrix, graph model, scoring formulas, and Codex critique into v2.
Deliverable: radargraph-world-algorithm-v2.html.
Required sections: one-sentence thesis; algorithm principles; graph schema; ranking formulas; run protocol; benchmark harness; implementation phases; safety gates; why it could be PageRank-level; why it could fail.
Stop gate: If the artifact reads like inspiration instead of a spec, it fails.
Hour 7: Benchmark harness design
Objective: Specify how to prove lift.
Fixture types: product differentiation fixture; fast-moving AI tooling fixture; local NYC/culture radar fixture; science learning/radar fixture; founder decision memo fixture; code/repo risk radar fixture.
Comparison arms: generic model single prompt; generic deep research prompt; Grok-only radar; Padawan single-pass synthesis; RadarGraph v1; RadarGraph with Codex critique.
Scoring: blind human preference; claim correctness; source novelty; source quality; actionability; critique survival; time/cost; reuse benefit on next run.
Deliverable: benchmark spec and sample scoring rubric.
Stop gate: If no blind or semi-blind comparison is possible, call that out.
Hour 8: Implementation architecture
Objective: Turn the algorithm into a buildable local system.
MVP components: radargraph/schema.py or JSON schema; radargraph/store.py backed by SQLite or JSONL first; radargraph/scoring.py for transparent formulas; radargraph/run_packet.py to store every run; radargraph/bench.py for fixture evaluation; radargraph/prompts/ for Grok/Codex/Padawan role prompts; radargraph/artifacts/ for generated reports.
Deliverable: repo-agnostic implementation plan; no code unless approved.
Stop gate: Keep implementation local-first; no infra, cloud, or credential dependencies.
Hour 9: Algorithm stress tests
Objective: Attack the algorithm with bad cases.
Stress cases: viral false claim with many citations; stale authoritative source; high-taste but low-evidence output; source monoculture; social hype around a weak product; adversary over-pruning a good insight; user preference drift; conflicting expert sources; missing primary data; expensive search path with low yield.
Deliverable: failure-mode table and required mitigations.
Stop gate: If the algorithm cannot say “not enough evidence,” it fails.
Hour 10: Product doctrine
Objective: Convert the algorithm into positioning and design constraints.
Doctrine draft: Padawan is not a chatbot. Padawan is not just personal memory. Padawan is not generic deep research. Padawan is a compounding source-intelligence layer for high-stakes judgment. The UI should expose provenance, deltas, critique, and graph memory — not hide them.
Deliverable: category narrative + anti-positioning.
Stop gate: If a competitor can copy the sentence without copying the machinery, rewrite it.
Hour 11: Final adversarial review
Objective: Decide whether the algorithm is truly differentiated or still pasta.
Review questions: What is the PageRank-like insight in one formula/paragraph? What is impossible to copy quickly? What depends on Hareesh’s taste vs system design? What evidence would prove this is not better? What should be built first? What should not be built? What is the smallest demo that makes the difference obvious?
Deliverable: final critique and patch list.
Stop gate: Do not publish final artifact until the top critique is addressed or explicitly accepted as residual risk.
Hour 12: Master artifact
Objective: Deliver the final 12h spec.
Deliverable: radargraph-world-algorithm-master-YYYYMMDD-HHMM.html.
Must include: executive thesis; PageRank analogy and difference; algorithm overview; graph schema; scoring formulas; pass protocol; source-radar policy; Codex adversarial policy; benchmark suite; safety model; implementation roadmap; what would change the plan; source trail; known blockers.
Stop gate: If tools failed, artifact must say exactly what failed and what fallback was used.
Concrete algorithm design target
Working name: RadarRank
RadarRank is the ranking layer inside RadarGraph.
The core ranking target is not webpages; it is decision-useful claims.
A claim’s value should rise when:
- multiple independent high-quality sources support it,
- it survives adversarial critique,
- it changes or improves a decision,
- it connects to durable graph knowledge,
- it has high novelty relative to generic baselines,
- and it has previously led to user-approved or benchmark-winning outputs.
A claim’s value should fall when:
- it is supported by source monoculture,
- it is contradicted by credible evidence,
- it is unfalsifiable,
- it is stale or context-misaligned,
- it is only stylistically pleasing,
- or it survives only because no adversary tested it.
Draft formula shape
ClaimRank(c) =
EvidenceSupport(c)
× SourceIndependence(c)
× CritiqueSurvival(c)
× DecisionLeverage(c)
× NoveltyAgainstBaseline(c)
× TasteGeneralization(c)
- ContradictionPenalty(c)
- StalenessPenalty(c)
- MonoculturePenalty(c)
- UnresolvedRiskPenalty(c)
Why this is PageRank-like
PageRank used the link graph to rank pages by structural endorsement, not only keyword match.
RadarRank should use the claim/source/critique/user-decision graph to rank research outputs by structural survivability and decision utility, not only model fluency.
Tests / validation
Sprint validation
- Each major artifact has source trail or explicit failure note.
- Each major algorithm claim has either source support or hypothesis label.
- Codex critique is incorporated or rejected with rationale.
- Final spec contains formulas, not just prose.
- Benchmark suite can be run manually on at least one fixture.
Future implementation validation
- Unit tests for graph schema serialization.
- Unit tests for scoring monotonicity:
- more independent support increases score,
- credible contradiction decreases score,
- unresolved high-severity critique decreases score,
- source monoculture penalty works,
- stale source penalty works.
- Golden fixture tests:
- RadarGraph output beats single-pass baseline on at least 4/6 quality axes.
- Regression tests:
- no final claims without evidence IDs,
- no hidden source-free confidence labels,
- no new search after final synthesis gate.
Operational safety controls
Cron controls
- Use at most one 12h sprint job.
- Do not let the job create/update/remove other jobs.
- Deliver only versioned artifacts or real blockers.
- Stop after 12 runs.
- Keep final synthesis separate from noisy interim runs if chat gets spammy.
Worker controls
- Codex runs in isolated scratch git repo or read-only prompt mode.
- No repo writes unless explicitly approved.
- Preserve Codex transcript/outputs before citing them.
- Kill worker if it asks for credentials, approvals, billing, or broad filesystem access.
Source controls
- Grok receipt path required when Grok succeeds.
- Failed Grok is a named failure, not silently substituted.
- Primary sources preferred for technical claims.
- Web snippets are leads, not proof.
Artifact controls
- HTML artifacts only under this profile’s brief-viewer path.
- No source/tool logs dumped into Telegram.
- Verification is internal; user sees the artifact pointer.
Risks, tradeoffs, and open questions
Risks
- Twelve hours may over-optimize a spec before real fixtures prove it. Mitigation: force benchmark fixture design early.
- The PageRank analogy may seduce us into graph math too soon. Mitigation: every score must map to an observable quality improvement.
- Codex may be slow or blocked. Mitigation: use it opportunistically for critique; do not block the entire sprint.
- Grok may fail or return transport errors. Mitigation: retry with receipts; fallback to web/source search and label degradation.
- Taste data may be sparse. Mitigation: use blind comparisons and explicit rejected-pattern capture.
Tradeoffs
- More automation increases search breadth but raises noise and safety risk.
- More graph structure improves compounding but can become bureaucracy.
- More formulas improve auditability but can create fake precision.
- More adversarial review improves rigor but can over-prune weird valuable ideas.
Open questions
- Should the first implementation live inside Hermes, a separate
radargraphrepo, or a Padawan vault prototype? - What are the first 6 benchmark fixtures Hareesh cares about most?
- Should user preference/taste data be explicit ratings, implicit edits, or both?
- What is the minimum demo that makes RadarGraph visibly better than generic deep research?
Approval checkpoint
Before execution, Hareesh should approve:
- Whether to pause/update the existing World Algorithm cron.
- Whether the 12h sprint should deliver hourly artifacts or only a final master artifact plus critical blockers.
- Whether Codex is allowed to write scratch files in this profile’s tmp directory.
- Whether the output should remain design-only or include a local prototype harness.
Recommended default:
- Pause noisy duplicate jobs if any.
- Run one 12h sprint job or manual supervised loop, not both.
- Deliver only the master artifact plus maybe 2–3 major upgrades.
- Allow scratch-only Codex files, no repo writes.
- End with spec + benchmark + prototype plan, not production code.