RadarGraph World Algorithm Safe Build Plan

Planning only. Do not implement this plan, modify product code, launch new crons, restart gateways, push commits, or let worker agents write external state until Hareesh approves the execution scope.

Goal: In 12 hours, produce a founder-grade, technically implementable “world algorithm” for Padawan RadarGraph: a source-intelligence ranking and synthesis system with PageRank-level ambition, using Grok radar, Codex adversarial/protocol engineering, and a compounding evidence graph.

Architecture: Treat this as a research-and-systems design sprint, not a hype sprint. The output is a versioned algorithm spec plus a minimal benchmark harness plan. The system’s core is a graph where sources, claims, critiques, searches, user judgments, and successful tactics become scored nodes/edges that improve future research trajectories.

Tech Stack: Hermes cron/session orchestration, Grok CLI for live radar, Codex CLI for adversarial/system critique, local HTML/Markdown artifacts, optional local SQLite/JSONL prototype spec, web/arXiv/source fetches, no production writes.

Useful anchors: PageRank paper, Codex CLI, Hermes Agent docs, arXiv.


Safety posture

Non-negotiable boundaries

Risk model

This sprint can fail in five main ways:

  1. Grandiosity drift: “beat every product” becomes abstract hype instead of measurable axes.
  2. Tool overreach: crons/workers multiply, become noisy, or touch unsafe surfaces.
  3. Source laundering: Grok or web snippets get promoted into facts without receipts.
  4. Algorithm theater: formulas look serious but do not improve outputs.
  5. Taste overfit: system learns Hareesh’s prose/taste but not durable research judgment.

The plan below is designed to force concrete outputs, measurable deltas, and stop gates.


Current context / assumptions

Known context

Assumptions


The central product thesis to test

RadarGraph should be to AI research agents what PageRank was to search: not merely a better model call, but a ranking/compounding structure that makes better information surface because the system represents relationships better than competitors do.

The candidate PageRank-like insight:

A research result is valuable when high-credibility sources, independent claims, adversarial critiques, and downstream user decisions mutually reinforce each other over time — and when the graph penalizes ungrounded fluency, source monoculture, unresolved critique, and non-actionable synthesis.

So the “world algorithm” must rank not only sources, but also claims, critique tickets, search trajectories, synthesis versions, user judgments, reusable tactics, and future radar priorities.


Proposed 12-hour sprint architecture

Sprint outputs

By the end of 12 hours, produce:

  1. Master algorithm spec as HTML + Markdown.
  2. Data model for the graph.
  3. Core scoring formulas with rationale and failure modes.
  4. Run protocol for Grok → Padawan → Codex → targeted radar → synthesis.
  5. Benchmark suite spec comparing against generic deep research products.
  6. Prototype build plan for a local harness.
  7. Safety/quality gates for future automation.
  8. Open questions / kill criteria that could prove the algorithm is not special.

Workstream roles


Safe execution timeline

Hour 0: Freeze and align

Objective: Prevent unsafe automation and clarify the sprint contract.

Actions:

  1. Inspect active crons.
  2. If a World Algorithm cron is active from the interrupted turn, either pause it until Hareesh approves this plan or update it to follow this plan exactly.
  3. Confirm no new cron should recursively create/modify crons.
  4. Set all sprint jobs to deliver artifacts only, not chat walls.
  5. Establish one scratch directory: /Users/openclaw/.hermes/profiles/personal/tmp/radargraph-world-algo-YYYYMMDD/.
  6. Establish one artifact directory: /Users/openclaw/.hermes/profiles/personal/brief-viewer/.

Deliverable: short Telegram checkpoint or artifact if long.

Stop gate: If automation state is ambiguous, pause the sprint jobs and proceed manually.

Hour 1: Frontier map

Objective: Build a source-backed map of relevant “latest advancements.”

Source lanes: graph ranking / PageRank successors; neural retrieval, reranking, hybrid search; knowledge graphs and claim graphs; agent trajectory evaluation; multi-agent debate, critique, verifier loops; source credibility and provenance; active learning / bandits for search policy; preference learning / taste calibration; deep research agent benchmarks and competitors.

Actions: Run Grok radar for frontier concepts and named systems; run web/source search for primary anchors; record each lead with source id, title, url, source class, claim relevance, reliability note, and accessed date.

Deliverable: frontier-map-v1.md or HTML section.

Stop gate: No source-backed algorithm claims may enter the master spec yet; this is only the map.

Hour 2: Competitor baseline and attack surface

Objective: Define what “beat them” means in concrete product terms.

Baselines: Generic ChatGPT/Claude/Gemini deep research; Perplexity-style answer engine; Manus/Genspark-style autonomous report agent; Notion/Granola/Limitless-style memory/workflow products; internal Padawan single-pass output.

Quality axes: novel source yield; claim grounding and traceability; adversarial survival; decision usefulness; taste/voice fit without mimicry; trajectory lift between passes; graph reuse in future runs; time/cost/noise efficiency.

Deliverable: baseline matrix.

Stop gate: If “beat them” cannot be measured on an axis, that axis is not allowed in the product claim.

Hour 3: Data model v1

Objective: Design the graph state that compounds.

Core nodes: SourceNode, AuthorNode, ClaimNode, QuestionNode, SearchTrajectoryNode, CritiqueNode, SynthesisNode, UserJudgmentNode, TacticNode, BenchmarkFixtureNode.

Core edges: SUPPORTS, CONTRADICTS, CITES, DERIVED_FROM, CRITIQUES, RESOLVES, FAILED_BY, PREFERRED_OVER, REUSED_IN, UPGRADED_BY.

Deliverable: schema spec with JSON examples.

Stop gate: If the model cannot represent why a future answer got better, it is not a compounding graph.

Hour 4: Ranking formula candidates

Objective: Draft PageRank-like scoring primitives.

Candidate scores:

  1. SourceRank: credibility × independence × recency × proximity to primary evidence × historical usefulness.
  2. ClaimRank: support strength × source independence × contradiction penalty × critique survival × decision leverage.
  3. CritiqueRank: severity × target load-bearingness × specificity × falsifiability × historical yield.
  4. TrajectoryRank: novelty yield × claim flips × critique resolution × time cost × downstream reuse.
  5. TasteFitScore: Hareesh preference signal × blind win rate × anti-slop penalty × non-mimicry constraint.
  6. ActionRank: expected founder decision impact × reversibility × cost × speed × evidence confidence.

Deliverable: formulas plus examples.

Stop gate: Every formula must name a gaming/failure mode.

Hour 5: Codex adversarial protocol audit

Objective: Treat the algorithm like software and attack it.

Codex prompt should ask: Where can this scoring system be gamed? Which variables are unobservable or fake precision? Which graph updates create feedback loops or taste overfit? What minimal tests would catch algorithm theater? What data model will break first? Which ranking formula has the highest risk of Goodharting?

Deliverable: Codex critique receipt + accepted/rejected findings.

Stop gate: Do not accept Codex critique wholesale. Padawan must classify each finding as accepted, partially accepted, rejected, or needs evidence.

Hour 6: RadarGraph v2 synthesis

Objective: Merge frontier map, baseline matrix, graph model, scoring formulas, and Codex critique into v2.

Deliverable: radargraph-world-algorithm-v2.html.

Required sections: one-sentence thesis; algorithm principles; graph schema; ranking formulas; run protocol; benchmark harness; implementation phases; safety gates; why it could be PageRank-level; why it could fail.

Stop gate: If the artifact reads like inspiration instead of a spec, it fails.

Hour 7: Benchmark harness design

Objective: Specify how to prove lift.

Fixture types: product differentiation fixture; fast-moving AI tooling fixture; local NYC/culture radar fixture; science learning/radar fixture; founder decision memo fixture; code/repo risk radar fixture.

Comparison arms: generic model single prompt; generic deep research prompt; Grok-only radar; Padawan single-pass synthesis; RadarGraph v1; RadarGraph with Codex critique.

Scoring: blind human preference; claim correctness; source novelty; source quality; actionability; critique survival; time/cost; reuse benefit on next run.

Deliverable: benchmark spec and sample scoring rubric.

Stop gate: If no blind or semi-blind comparison is possible, call that out.

Hour 8: Implementation architecture

Objective: Turn the algorithm into a buildable local system.

MVP components: radargraph/schema.py or JSON schema; radargraph/store.py backed by SQLite or JSONL first; radargraph/scoring.py for transparent formulas; radargraph/run_packet.py to store every run; radargraph/bench.py for fixture evaluation; radargraph/prompts/ for Grok/Codex/Padawan role prompts; radargraph/artifacts/ for generated reports.

Deliverable: repo-agnostic implementation plan; no code unless approved.

Stop gate: Keep implementation local-first; no infra, cloud, or credential dependencies.

Hour 9: Algorithm stress tests

Objective: Attack the algorithm with bad cases.

Stress cases: viral false claim with many citations; stale authoritative source; high-taste but low-evidence output; source monoculture; social hype around a weak product; adversary over-pruning a good insight; user preference drift; conflicting expert sources; missing primary data; expensive search path with low yield.

Deliverable: failure-mode table and required mitigations.

Stop gate: If the algorithm cannot say “not enough evidence,” it fails.

Hour 10: Product doctrine

Objective: Convert the algorithm into positioning and design constraints.

Doctrine draft: Padawan is not a chatbot. Padawan is not just personal memory. Padawan is not generic deep research. Padawan is a compounding source-intelligence layer for high-stakes judgment. The UI should expose provenance, deltas, critique, and graph memory — not hide them.

Deliverable: category narrative + anti-positioning.

Stop gate: If a competitor can copy the sentence without copying the machinery, rewrite it.

Hour 11: Final adversarial review

Objective: Decide whether the algorithm is truly differentiated or still pasta.

Review questions: What is the PageRank-like insight in one formula/paragraph? What is impossible to copy quickly? What depends on Hareesh’s taste vs system design? What evidence would prove this is not better? What should be built first? What should not be built? What is the smallest demo that makes the difference obvious?

Deliverable: final critique and patch list.

Stop gate: Do not publish final artifact until the top critique is addressed or explicitly accepted as residual risk.

Hour 12: Master artifact

Objective: Deliver the final 12h spec.

Deliverable: radargraph-world-algorithm-master-YYYYMMDD-HHMM.html.

Must include: executive thesis; PageRank analogy and difference; algorithm overview; graph schema; scoring formulas; pass protocol; source-radar policy; Codex adversarial policy; benchmark suite; safety model; implementation roadmap; what would change the plan; source trail; known blockers.

Stop gate: If tools failed, artifact must say exactly what failed and what fallback was used.


Concrete algorithm design target

Working name: RadarRank

RadarRank is the ranking layer inside RadarGraph.

The core ranking target is not webpages; it is decision-useful claims.

A claim’s value should rise when:

A claim’s value should fall when:

Draft formula shape

ClaimRank(c) =
  EvidenceSupport(c)
  × SourceIndependence(c)
  × CritiqueSurvival(c)
  × DecisionLeverage(c)
  × NoveltyAgainstBaseline(c)
  × TasteGeneralization(c)
  - ContradictionPenalty(c)
  - StalenessPenalty(c)
  - MonoculturePenalty(c)
  - UnresolvedRiskPenalty(c)

Why this is PageRank-like

PageRank used the link graph to rank pages by structural endorsement, not only keyword match.

RadarRank should use the claim/source/critique/user-decision graph to rank research outputs by structural survivability and decision utility, not only model fluency.


Tests / validation

Sprint validation

Future implementation validation


Operational safety controls

Cron controls

Worker controls

Source controls

Artifact controls


Risks, tradeoffs, and open questions

Risks

Tradeoffs

Open questions


Approval checkpoint

Before execution, Hareesh should approve:

  1. Whether to pause/update the existing World Algorithm cron.
  2. Whether the 12h sprint should deliver hourly artifacts or only a final master artifact plus critical blockers.
  3. Whether Codex is allowed to write scratch files in this profile’s tmp directory.
  4. Whether the output should remain design-only or include a local prototype harness.

Recommended default: