The Lookout
The Lookout
A curated reading room: the primary sources, courses, and papers worth a practitioner's hours, each with two lines of our own on what it says and what it changes for a practice. Compiled October 3, 2026; refreshed monthly with the Letters. Every link is to the publisher's own page; nothing here is ours except the commentary. Links are verified at each site build; an entry whose link fails is removed, not left to rot.
Plate I · the lookout · Stoa MMXXVIThe round watchtower at the corner of the wall and a dark bronze telescope on its tripod facing the sea, a gold ring on the tube, the sun from the left.
I. How to think about agents and harnesses
- Building Effective Agents (Anthropic, December 2024)
What it says: most production "agents" are simple composable workflows, and the simplest architecture that works beats the elaborate one. What it changes: the catalog's design rule in one essay; read it before buying any orchestration product.
- Claude documentation: prompting and evaluation guides (Anthropic)
What it says: concrete rubrics, reason-then-verdict grading, and prompt structure, from the publisher. What it changes: the Grading Room's method is drawn from here; every evaluator rubric should be checked against it.
II. The papers behind the instruments
- Attention Is All You Need (Vaswani et al., 2017)
What it says: the transformer architecture every current model descends from. What it changes: nothing operational, and everything conceptual; the one paper worth reading to stop treating the models as magic.
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Zheng et al., 2023)
What it says: models grading models show position, verbosity, and self-enhancement biases, measured. What it changes: why the Grading Room uses a separate room, binary checks, and a human validation loop instead of a score.
- Lost in the Middle: How Language Models Use Long Contexts (Liu et al., 2023)
What it says: models attend best to the start and end of a long context and worst to the middle. What it changes: where the rules go in an instruction file (top and bottom), and why the catalog keeps instruction files short.
- ReAct: Synergizing Reasoning and Acting in Language Models (Yao et al., 2022)
What it says: interleaving reasoning with tool actions outperforms either alone. What it changes: the shape of every scheduled loop in the Loop Library; the model reasons, acts, and reports, in that order.
- Constitutional AI: Harmlessness from AI Feedback (Bai et al., 2022)
What it says: a written set of principles can steer model behavior through self-critique. What it changes: the reason an instruction framework works at all; principles written down compound.
- Self-Preference Bias in LLM-as-a-Judge (Wataoka et al., 2024)
What it says: judges favor output that resembles their own, and the effect tracks perplexity. What it changes: the independence rule, stated with a number behind it.
III. Courses and training from the labs and the clouds
- Anthropic Academy / Claude courses
What it says: structured courses on prompting, tool use, and building with Claude. What it changes: the self-serve path for a team member who needs the fundamentals before an install; assign before the first session.
- OpenAI Academy
What it says: free foundational and role-based learning for ChatGPT and the API. What it changes: the equivalent assignment for ChatGPT-first teams.
- Google Cloud Skills Boost, generative AI learning path
What it says: vendor training on generative AI concepts and Google's tooling. What it changes: useful for operators on Google Workspace; the concepts transfer, the tooling is theirs.
- Microsoft Learn: AI and Copilot training
What it says: Microsoft's structured training for Copilot and Azure AI. What it changes: the reference for operators standardized on Microsoft 365, where the catalog's compressed path meets their admin's controls.
- AWS Skill Builder, generative AI courses
What it says: Amazon's foundational AI training. What it changes: relevant to operators whose developers host client-owned builds on AWS; not a buyer-facing path.
- NVIDIA Deep Learning Institute
What it says: technical courses from the hardware layer up. What it changes: for the operator's engineer, not the operator; listed so the engineer has a vendor-neutral starting point.
- DeepLearning.AI short courses
What it says: one-hour courses on prompting, agents, evaluation, and retrieval, many co-taught with the labs. What it changes: the fastest honest way to learn what "evals" and "RAG" mean before someone sells you either.
IV. Standards and guardrails
- NIST AI Risk Management Framework
What it says: the public standard for governing AI risk: map, measure, manage, govern. What it changes: the vocabulary an operator's counsel and insurer will use; the Compliance Pack's screens map onto it.
- FTC: AI and consumer protection guidance
What it says: the regulator's running commentary on deceptive AI claims, fake reviews, and endorsements. What it changes: why no catalog product promises an outcome and why the Review Response System screens for incentivized reviews.
- HUD: Fair Housing Act resources
What it says: the statute and the agency's guidance, including on advertising and algorithms. What it changes: the Fair Housing Sweep's authority; read the guidance, not the summaries.
V. Reading the market
- Freddie Mac Primary Mortgage Market Survey
What it says: the weekly rate everyone quotes. What it changes: The Record cites it every month; know its lag against the daily indices.
- Zillow Research data
What it says: public monthly series for values, inventory, and listings by metro and county. What it changes: the series behind most of The Record's charts; the terms require attribution, which The Record gives.
- FRED (Federal Reserve Bank of St. Louis)
What it says: the public macro and housing series, including the Atlanta Case-Shiller index. What it changes: the verification layer for every number in The Record.
Educational, not professional advice. Nothing on this page is affiliated with or endorsed by its publisher.
