Skip to content

Support authenticated ProofLoop production rooms#204

Merged
HomenShum merged 1 commit into
mainfrom
codex/nodeagent-prod-route-evidence
Jul 16, 2026
Merged

Support authenticated ProofLoop production rooms#204
HomenShum merged 1 commit into
mainfrom
codex/nodeagent-prod-route-evidence

Conversation

@HomenShum

Copy link
Copy Markdown
Owner

Summary

  • accept explicit per-adapter fresh room URLs for external adapter proofs
  • accept an explicit Playwright storage-state fixture through PROOFLOOP_AUTH_STORAGE_STATE
  • join and create the blank sheet through ordinary production UI
  • keep credentials and browser session data out of source and receipts

Verification

  • focused auth-bootstrap and benchmark-board suites: 10/10
  • npm run floor: 350 files, 2,487 tests

Honest boundary

Production currently uses GitHub-only auth. No approved Playwright storage-state fixture exists in this shell, so unattended isolated-browser proof remains blocked until a user-generated fixture is supplied.

@vercel

vercel Bot commented Jul 16, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
noderoom Building Building Preview, Comment Jul 16, 2026 9:04pm

Request Review

@HomenShum
HomenShum merged commit dd029c6 into main Jul 16, 2026
8 of 10 checks passed
@github-actions

Copy link
Copy Markdown

Scaffold Handoff — For Your Coding Agent

Your coding agent (Codex, Claude Code, etc.) should apply the accepted
scaffold proposals below. Do NOT touch any immutable files.

Immutability Check

Mode: advisory

✅ No immutable files were modified in this branch.

Changed Files

  • docs/eval/OFFICIAL_BENCHMARK_READINESS.md
  • docs/eval/OFFICIAL_BENCHMARK_TASK_COVERAGE.md
  • docs/eval/OPENROUTER_CONVEX_BENCHMARK.md
  • docs/eval/agent-improvement-loop.md
  • docs/eval/agent-improvement-loop.svg
  • docs/eval/agent-improvement-loop/20260716T210525Z.json
  • docs/eval/agent-improvement-loop/latest.json
  • docs/eval/agent-workspace-sandbox-smoke.json
  • docs/eval/algorithm-artifact-smoke.json
  • docs/eval/bankertoolbench-official-contract.json
  • docs/eval/docker-sandbox-probe.json
  • docs/eval/eval-runs.jsonl
  • docs/eval/halo-convex-context-telemetry.json
  • docs/eval/halo-self-improvement-smoke.json
  • docs/eval/halo-variant-selection.json
  • docs/eval/official-benchmark-readiness.json
  • docs/eval/official-benchmark-task-coverage.json
  • docs/eval/openrouter-convex-benchmark.json
  • docs/eval/professional-catalog-proofs.json
  • docs/eval/professional-proof-ledger.json
  • docs/eval/spreadsheetbench-chart-visual-probe.json
  • docs/eval/traces/credit/20260716T210534319Z-4a3d5a5a_dirty.1fe8b6081dbca29a/cascade-healthy.json
  • docs/eval/traces/credit/20260716T210534319Z-4a3d5a5a_dirty.1fe8b6081dbca29a/delta-incomplete.json
  • docs/eval/traces/credit/20260716T210534319Z-4a3d5a5a_dirty.1fe8b6081dbca29a/mapping-correct.json
  • docs/eval/traces/credit/20260716T210534319Z-4a3d5a5a_dirty.1fe8b6081dbca29a/mapping-misbind.json
  • docs/eval/traces/credit/20260716T210534319Z-4a3d5a5a_dirty.1fe8b6081dbca29a/summit-stressed.json
  • docs/eval/traces/ladder/20260716T210533863Z-4a3d5a5a_dirty.5545cd85b13f65fa/ladder_L1_read_scripted.json
  • docs/eval/traces/ladder/20260716T210533863Z-4a3d5a5a_dirty.5545cd85b13f65fa/ladder_L2_edit_scripted.json
  • docs/eval/traces/ladder/20260716T210533863Z-4a3d5a5a_dirty.5545cd85b13f65fa/ladder_L3_conflict_scripted.json
  • docs/eval/traces/ladder/20260716T210533863Z-4a3d5a5a_dirty.5545cd85b13f65fa/ladder_L4_blocked_scripted.json
  • docs/eval/traces/ladder/20260716T210533863Z-4a3d5a5a_dirty.5545cd85b13f65fa/ladder_L5_large_range_scripted.json
  • docs/eval/traces/ladder/20260716T210533863Z-4a3d5a5a_dirty.5545cd85b13f65fa/ladder_L6_long_horizon_scripted.json
  • docs/eval/traces/ladder/20260716T210533863Z-4a3d5a5a_dirty.5545cd85b13f65fa/ladder_L7_resume_scripted.json

Needs Adversarial Review — Do NOT Apply Yet

These proposals passed the reject check but have not been approved by
an adversarial reviewer. A human or frozen LLM judge must approve them first.

  • scaf-001 (AGENTS.md): Add explicit instruction for step spreadsheetbench-runner-fixture: Step spreadsheetbench-runner-fixture failed — scaffold may need explicit instruction or evidence assertion.
  • scaf-002 (AGENTS.md): Add explicit instruction for step convex-boundaries: Step convex-boundaries failed — scaffold may need explicit instruction or evidence assertion.

Safety Boundary

Agent may improve the scaffold.
Agent may NOT weaken the proof gate.

Immutable files (never modify):

  • scripts/proofloop.mjs
  • scripts/agent-improvement-loop.ts
  • tests/harnessChangeEval.test.ts
  • .github/workflows/
  • src/eval/evalTrustPolicy.ts
  • src/eval/architectureBudget.ts
  • evals/evalStore.ts

Scaffold files (safe to modify):

  • AGENTS.md
  • CLAUDE.md
  • proofloop/scenarios/*.yaml
  • proofloop/rubrics/*.yaml
  • proofloop/subagents/*.md
  • proofloop/adapters/*.js
  • .proofloop/memory.jsonl
  • src/nodeagent/models/prompts/systemPrompt.ts

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant