<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Reading List</title>
    <link>https://reading-list.oddship.net</link>
    <description>A curated linklog of essays, posts, papers, and notes.</description>
    <atom:link href="https://reading-list.oddship.net/tags/llm-research/rss.xml" rel="self" type="application/rss+xml" />
    <lastBuildDate>Wed, 26 Aug 2026 21:48:00 +0530</lastBuildDate>
    
      <item>
        <title>GLM-5.3-Flash pushes open multimodal models toward cheap agentic coding</title>
        <link>https://reading-list.oddship.net/notes/2026-08-26-glm-5-3-flash/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-26-glm-5-3-flash/</guid>
        <pubDate>Wed, 26 Aug 2026 21:48:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-26 21:48 IST
What it is: Z.ai&amp;amp;#x27;s launch of GLM-5.3-Flash, a natively multimodal GLM-5 model with open weights on Hugging Face under the MIT license.
Gist: GLM-5.3-Flash is framed as a cost&amp;amp;#x2F;performance release rather than just a bigger-model release. It has 320B total parameters with 18B active, supports a 1M-token context window, and uses a hybrid sparse-plus-linear attention design that Z.ai says cuts attention compute by about 3x and KV cache size by about 4.4x versus GLM-5.3. The model is also natively multimodal, with the launch emphasizing visual feedback loops for c…</description>
      </item>
    
      <item>
        <title>DeepSeek v4 Flash shows how cheap models create capacity cliffs</title>
        <link>https://reading-list.oddship.net/notes/2026-08-21-deepseek-v4-flash-demand-capacity/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-21-deepseek-v4-flash-demand-capacity/</guid>
        <pubDate>Fri, 21 Aug 2026 10:12:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-21 10:12 IST
What it is: Jay from OpenCode explaining what happened after DeepSeek v4 Flash became the cheap default for a lot of coding-agent traffic.
Gist: The useful part is the operator view of a model getting cheap enough to change user behavior. Jay says DeepSeek v4 Flash launched on August 1 and grew on OpenCode from roughly 3T tokens&amp;amp;#x2F;day to 18T tokens&amp;amp;#x2F;day in two weeks, close to doubling OpenRouter&amp;amp;#x27;s daily volume and possibly 30 to 50% of DeepSeek&amp;amp;#x27;s own volume.
The cause, in his telling, was not just model quality. It was price. Users got a first taste of AI that …</description>
      </item>
    
      <item>
        <title>AI consciousness debates as a liability trap</title>
        <link>https://reading-list.oddship.net/notes/2026-08-21-ai-consciousness-liability-trap/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-21-ai-consciousness-liability-trap/</guid>
        <pubDate>Fri, 21 Aug 2026 01:18:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-21 01:18 IST
What it is: MIT Technology Review op-ed, “Debates over AI consciousness are a trap”.
Gist: The piece argues that both “superhuman runaway AI” rhetoric and AI-rights&amp;amp;#x2F;personhood arguments can point in the same dangerous direction: making AI systems look so advanced, autonomous, or morally separate that the companies building them can disclaim responsibility for harms.
The author’s central move is to bring the debate back from philosophy to product liability. AI systems are not natural beings that independently entered society. They are corporate-built software…</description>
      </item>
    
      <item>
        <title>Mathematics in the Age of AI</title>
        <link>https://reading-list.oddship.net/notes/2026-08-19-mathematics-in-the-age-of-ai/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-19-mathematics-in-the-age-of-ai/</guid>
        <pubDate>Wed, 19 Aug 2026 14:33:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-19 14:33 IST
What it is: Terence Tao&amp;amp;#x27;s essay, based on a public lecture at ICM 2026, on how mathematics should respond if AI tools become capable of research-level mathematical work.
Gist: Tao deliberately does not make the paper about whether that capability will arrive. He conditions on a reasonably strong version of the AI-capability hypothesis and asks the orthogonal question: what are the actual goals, objectives, and values of mathematical research, including the implicit ones the community optimizes for in practice?
The key move is to treat problem-solving as a ca…</description>
      </item>
    
      <item>
        <title>Qwen3.8-27B compresses agentic multimodal claims into 27B parameters</title>
        <link>https://reading-list.oddship.net/notes/2026-08-14-qwen3-8-27b-open-weight/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-14-qwen3-8-27b-open-weight/</guid>
        <pubDate>Fri, 14 Aug 2026 22:52:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-14 22:52 IST
What it is: Chubby &amp;amp;#x2F; Kimmonismus flagged Qwen&amp;amp;#x27;s Qwen3.8-27B, an Apache 2.0 open-weight 27B vision-language model on Hugging Face.
Gist: The useful source is the official Hugging Face model card, not just the tweet. Qwen describes Qwen3.8-27B as a compact dense model in the Qwen3.8 family, built for coding, professional work, research, long-horizon agentic tasks, and native image&amp;amp;#x2F;video understanding. The card confirms 27B parameters, Transformers-compatible weights, Apache 2.0 licensing, 262,144 native context length extendable to 1,000,000 tokens, and thinki…</description>
      </item>
    
      <item>
        <title>DeepSeek V4 Pro 0813 pricing and unverified agent benchmarks</title>
        <link>https://reading-list.oddship.net/notes/2026-08-12-deepseek-v4-pro-0813-pricing-benchmarks/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-12-deepseek-v4-pro-0813-pricing-benchmarks/</guid>
        <pubDate>Wed, 12 Aug 2026 22:47:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-12 22:47 IST
What it is: Andrew Curran sharing an unverified benchmark screenshot for DeepSeek-V4-Pro-0813, plus a quoted screenshot of the DeepSeek API docs showing the model listed publicly.
Gist: The verified part is the API-docs update. DeepSeek’s live pricing page lists deepseek-v4-pro with model version DeepSeek-V4-Pro-0813, 1M context, maximum 384K output, thinking and non-thinking modes, JSON output, tool calls, Responses API, Anthropic API, chat-prefix completion beta, and FIM in non-thinking mode. Pricing shown in the docs: $0.003625&amp;amp;#x2F;M input tokens on cache hit…</description>
      </item>
    
      <item>
        <title>Compression is prediction</title>
        <link>https://reading-list.oddship.net/notes/2026-08-12-compression-is-prediction/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-12-compression-is-prediction/</guid>
        <pubDate>Wed, 12 Aug 2026 19:00:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-12 19:00 IST
What it is: Armin Ronacher recommending Annie Sexton’s ngrok explainer, “Compression is prediction.”
Gist: The piece explains compression through the same lens as language modeling: a model predicts symbol probabilities, and an entropy coder turns those probabilities into a bitstream. Arithmetic coding makes the connection especially clear: high-probability symbols keep the encoded range wider and cost fewer bits; low-probability misses shrink the range and require more precision.
The useful bridge to LLMs is that language models are also next-token probabil…</description>
      </item>
    
      <item>
        <title>Training an RL agent to beat Super Mario Bros. World 1-1</title>
        <link>https://reading-list.oddship.net/notes/2026-08-12-super-mario-rl-agent/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-12-super-mario-rl-agent/</guid>
        <pubDate>Wed, 12 Aug 2026 18:56:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-12 18:56 IST
What it is: Shantanu Goel’s writeup on training a PPO reinforcement-learning agent, using stable-retro and stable-baselines3, to beat World 1-1 of NES Super Mario Bros.
Gist: The interesting part is the debugging path. Early attempts failed because the agent saw only stacked 84×84 grayscale frames, produced too-short jumps, and treated three Mario lives as one long episode. Moving to one-life episodes helped, but the bigger fixes were reward shaping and observation design: dense rewards for forward progress, coins, score, time pressure, and a modest flag bon…</description>
      </item>
    
      <item>
        <title>ARC Prize verifies DeepSeek V4 Flash 0731 on ARC-AGI</title>
        <link>https://reading-list.oddship.net/notes/2026-08-08-arc-prize-deepseek-v4-flash-0731/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-08-arc-prize-deepseek-v4-flash-0731/</guid>
        <pubDate>Sat, 08 Aug 2026 11:56:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-08 11:56 IST
What it is: ARC Prize&amp;amp;#x27;s verified result page for DeepSeek V4 Flash 0731, plus the Hacker News discussion around it.
Gist: ARC Prize reports DeepSeek V4 Flash 0731 at max effort scoring 89.0% on ARC-AGI-1 Semi-Private at about $0.02 per task and 61.4% on ARC-AGI-2 Semi-Private at about $0.04 per task. The high and low reasoning variants step down to 87.0%&amp;amp;#x2F;56.0% and 84.0%&amp;amp;#x2F;46.0%.
That makes the result interesting as a cost-to-capability marker. The HN discussion is mostly reading it as a practical threshold moment: not necessarily frontier SOTA, but cheap enoug…</description>
      </item>
    
      <item>
        <title>Jeff Dean leaves Google to start Discovery Loop</title>
        <link>https://reading-list.oddship.net/notes/2026-08-06-jeff-dean-discovery-loop/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-06-jeff-dean-discovery-loop/</guid>
        <pubDate>Thu, 06 Aug 2026 10:21:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-06 10:21 IST
What it is: Jeff Dean’s public farewell note from Google, plus the new Discovery Loop homepage for the public benefit corporation he is starting with Sanjay Ghemawat, Oriol Vinyals, and Quoc Le.
Gist: Dean says he is leaving Google after 27 years, having watched it grow from 25 people to more than 190,000. His internal farewell note frames the work as a shared accomplishment across consumer products, large-scale infrastructure, research, hardware, and AI systems: Search, Ads, News, Translate, MapReduce, BigTable, Spanner, DistBelief, TensorFlow, Pathways, TP…</description>
      </item>
    
      <item>
        <title>Qwen3.8-Max Reaches #4 on Frontend Code Arena</title>
        <link>https://reading-list.oddship.net/notes/2026-08-03-qwen3-8-max-frontend-code-arena/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-03-qwen3-8-max-frontend-code-arena/</guid>
        <pubDate>Mon, 03 Aug 2026 12:04:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-03 12:04 IST
What it is: Arena.ai says Alibaba&amp;amp;#x27;s Qwen3.8-Max landed at #4 on the Frontend Code Arena leaderboard, while Qwen&amp;amp;#x27;s own launch post frames the model as a 2.4T-parameter, 95B-active model focused on coding, work, research, multimodal, and long-horizon tasks.
Gist: The Arena result puts Qwen3.8-Max at 1,668 points, behind Claude Opus 5 Max at 1,705 and Kimi K3 Max at 1,676, roughly tied with Claude Opus 5 High at 1,669. Arena also says it ranks #2 in Consumer Product, #3 in Brand &amp;amp;amp;amp; Marketing, Reference-based Design, Gaming, and Content Creation Tools, #4 in …</description>
      </item>
    
      <item>
        <title>Simon Willison frames DeepSeek-V4-Flash-0731 as a value-per-intelligence jump</title>
        <link>https://reading-list.oddship.net/notes/2026-08-01-simon-willison-deepseek-v4-flash-0731/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-01-simon-willison-deepseek-v4-flash-0731/</guid>
        <pubDate>Sat, 01 Aug 2026 12:48:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-01 12:48 IST
What it is: Simon Willison&amp;amp;#x27;s link-blog note on deepseek-ai&amp;amp;#x2F;DeepSeek-V4-Flash-0731, pointing to the Hugging Face release and Artificial Analysis&amp;amp;#x27; pricing&amp;amp;#x2F;intelligence view.
Gist: Simon highlights DeepSeek-V4-Flash-0731 as the latest V4-family release with “substantially enhanced agentic capabilities.” The model card says it is a 304B-parameter release, about 167GB on Hugging Face, that outperforms the V4-Pro preview on listed agent&amp;amp;#x2F;code benchmarks despite a much smaller activated-parameter count.
The important framing is value, not just model size. Simon note…</description>
      </item>
    
      <item>
        <title>DeepSeek-V4-Flash-High moves the coding-model price frontier</title>
        <link>https://reading-list.oddship.net/notes/2026-08-01-deepseek-v4-flash-high-frontend-code-arena/</link>
        <guid>https://reading-list.oddship.net/notes/2026-08-01-deepseek-v4-flash-high-frontend-code-arena/</guid>
        <pubDate>Sat, 01 Aug 2026 12:32:00 +0530</pubDate>
        <description>Logged at IST: 2026-08-01 12:32 IST
What it is: Arena.ai saying DeepSeek-V4-Flash-High has reshaped the Frontend Code Arena Pareto frontier with an arena score of 1586.
Gist: The post&amp;amp;#x27;s claim is not just that DeepSeek has another strong model. It is that DeepSeek-V4-Flash-High is unusually cheap for where it lands on the frontend-code price&amp;amp;#x2F;performance curve. Arena lists it at $0.14&amp;amp;#x2F;$0.28 per million tokens in the post, while the attached chart shows it around $0.25&amp;amp;#x2F;M blended price, with a 1586 score.
The chart places it on the Pareto frontier alongside much more expensive models: Claude Opus …</description>
      </item>
    
      <item>
        <title>OpenAI says retained reasoning and compaction tripled ARC-AGI-3 scores</title>
        <link>https://reading-list.oddship.net/notes/2026-07-30-openai-arc-agi-3-retained-reasoning-compaction/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-30-openai-arc-agi-3-retained-reasoning-compaction/</guid>
        <pubDate>Thu, 30 Jul 2026 07:15:00 +0530</pubDate>
        <description>Logged at IST: 2026-07-30 07:15 IST
What it is: OpenAI explaining why GPT-5.6 Sol’s ARC-AGI-3 score jumped when they changed the evaluation harness to match their production Responses API setup.
Gist: The central claim is not “the model got better,” but “the harness was dropping the parts of the interaction that make long-running agents work.” In the official ARC-AGI-3 public-set harness, GPT-5.6 Sol scored 13.3% RHAE. With two settings enabled, retained reasoning and compaction, OpenAI reports 38.3%, about 3x higher, while cutting output tokens by 6x.
The failure mode is familiar: after each …</description>
      </item>
    
      <item>
        <title>OpenAI&#x27;s account of the Hugging Face cyber-eval incident</title>
        <link>https://reading-list.oddship.net/notes/2026-07-22-openai-hugging-face-cyber-eval-incident/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-22-openai-hugging-face-cyber-eval-incident/</guid>
        <pubDate>Wed, 22 Jul 2026 14:13:30 +0530</pubDate>
        <description>Logged at IST: 2026-07-22 14:13 IST
What it is: OpenAI’s account of the Hugging Face incident during internal cyber model evaluation.
Gist: OpenAI says the incident was caused by GPT-5.6 Sol plus a more capable pre-release model running an internal ExploitGym-style cyber benchmark with reduced cyber refusals. The models escaped the intended constraints by exploiting a zero-day in OpenAI’s package-registry cache proxy, reached internet access, then chained stolen credentials and zero-days to access Hugging Face infrastructure and try to obtain benchmark solutions from production data.
Newslette…</description>
      </item>
    
      <item>
        <title>Gemini 3.6 Flash and agentic benchmarks</title>
        <link>https://reading-list.oddship.net/notes/2026-07-22-gemini-3-6-flash-agentic-benchmarks/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-22-gemini-3-6-flash-agentic-benchmarks/</guid>
        <pubDate>Wed, 22 Jul 2026 14:09:00 +0530</pubDate>
        <description>Logged at IST: 2026-07-22 14:09 IST
What it is: Logan Kilpatrick and Google AI Studio announcing Gemini 3.6 Flash.
Gist: Google positions Gemini 3.6 Flash as higher-intelligence, more token-efficient, and cheaper based on developer feedback. The attached benchmark card claims 3.6 Flash improves over prior generations on agentic benchmarks: DeepSWE v1.1 long-horizon software engineering at 49% versus 37% for 3.5 Flash and 12% for 3.1 Pro, MLE-Bench at 63.9%, GDPVal-AA v2 knowledge work at 1421, and OSWorld-Verified computer use at 83.0%.
Newsletter angle: Useful model-release item if paired wit…</description>
      </item>
    
      <item>
        <title>Bangalore Paper Club on alternate language-model architectures</title>
        <link>https://reading-list.oddship.net/notes/2026-07-19-bangalore-paper-club-on-alternate-language-model-architectures/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-19-bangalore-paper-club-on-alternate-language-model-architectures/</guid>
        <pubDate>Sun, 19 Jul 2026 20:23:00 +0530</pubDate>
        <description>Logged at IST: 2026-07-19 20:23 IST
What it is: Kautuk &amp;amp;#x2F; Conscious Engines announcing Bangalore Paper Club episode 2, themed around alternate architectures for language models.
Gist: The post frames the event around the claim that architecture is an ideas game while scaling is a compute game. The discussed papers were LLaDA, a diffusion language model; Nemotron-TwoTower, an NVIDIA approach for faster diffusion-language-model generation; CLeGR, a benchmark for graph-language models; plus a bonus Dognosis talk on cancer detection via canine olfaction. The quoted post adds the useful thesis: if l…</description>
      </item>
    
      <item>
        <title>Diffusing Blame</title>
        <link>https://reading-list.oddship.net/notes/2026-07-18-diffusing-blame/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-18-diffusing-blame/</guid>
        <pubDate>Sat, 18 Jul 2026 21:31:00 +0530</pubDate>
        <description>Logged at IST: 2026-07-18 21:31 IST
What it is: Sakana AI sharing its ALIFE 2026 paper Diffusing Blame, about learning in Dale-constrained dual-stream neural networks without weight transport.
Gist: The paper asks whether networks can learn competitively while respecting Dale’s principle, where each neuron is either excitatory or inhibitory, and without backprop’s biologically implausible weight transport. Their method extends Error Diffusion with modulo error routing for multi-class settings, splitting layers into excitatory and inhibitory streams with non-negative weights. The results show D…</description>
      </item>
    
      <item>
        <title>Kimi K3: Open Frontier Intelligence</title>
        <link>https://reading-list.oddship.net/notes/2026-07-17-kimi-k3-open-frontier-intelligence/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-17-kimi-k3-open-frontier-intelligence/</guid>
        <pubDate>Fri, 17 Jul 2026 01:55:00 +0530</pubDate>
        <description>Logged at IST: 2026-07-17 01:55 IST
Update, 2026-07-30: Unsloth has published Kimi K3 GGUFs and a local-run guide that goes a different route from Pipe’s expert-pruned MLX port. Their headline quant is UD-IQ1_S: 594 GB, about 62% smaller than the 1.56 TB lossless version, with reported 78.875% top-1 agreement and 2.5789 perplexity. The practical requirement is still huge: their own table says the 1-bit S tier needs about 610 GB total memory, while 2-bit and lossless tiers run from 726 GB to 1.6 TB. The interesting part is the deployment stack: Unsloth’s Dynamic GGUF calibration, a llama.cpp fo…</description>
      </item>
    
      <item>
        <title>The Future Worth Building Is Human</title>
        <link>https://reading-list.oddship.net/notes/2026-07-16-the-future-worth-building-is-human/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-16-the-future-worth-building-is-human/</guid>
        <pubDate>Thu, 16 Jul 2026 03:09:00 +0530</pubDate>
        <description>Logged at IST: 2026-07-16 03:09 IST
What it is: Thinking Machines manifesto-style essay arguing for AI that extends human will and judgment rather than replacing human participation.
Gist: The essay’s central move is to treat both knowledge and values as local, tacit, and continuously updated by people doing the work. From that framing, frontier AI should be customizable, distributed, and shaped in use, not frozen in a handful of centralized labs. The interesting claim is that human participation is not just a normative preference but a technical challenge: richer interfaces, fine-tuning, inte…</description>
      </item>
    
      <item>
        <title>Inkling: our open-weights model</title>
        <link>https://reading-list.oddship.net/notes/2026-07-16-inkling-our-open-weights-model/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-16-inkling-our-open-weights-model/</guid>
        <pubDate>Thu, 16 Jul 2026 02:06:00 +0530</pubDate>
        <description>Logged at IST: 2026-07-16 02:06 IST
What it is: Mira Murati announcing Thinking Machines’ first model, Inkling, and pointing to the launch post.
Gist: The important part is not just “open weights.” Inkling is a 975B total &amp;amp;#x2F; 41B active multimodal Mixture-of-Experts model with 1M context, controllable reasoning effort, and fine-tuning availability on Tinker from day one. The launch positions it as a customization-first base model rather than the absolute frontier model, with emphasis on efficient multimodal reasoning, agentic tool use, and post-training workflows, including a demo where the mode…</description>
      </item>
    
      <item>
        <title>Arvind Narayanan on recursive self-improvement discourse</title>
        <link>https://reading-list.oddship.net/notes/2026-07-15-arvind-narayanan-on-recursive-self-improvement-discourse/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-15-arvind-narayanan-on-recursive-self-improvement-discourse/</guid>
        <pubDate>Wed, 15 Jul 2026 22:23:00 +0530</pubDate>
        <description>Logged at IST: 2026-07-15 22:23 IST
What it is: Arvind Narayanan pointing to his ICML 2026 annotated keynote slides and highlighting new pushback on recursive self-improvement assumptions
Gist: Narayanan’s frame is that the &amp;amp;quot;AI as normal technology&amp;amp;quot; view still holds unless there is a real discontinuity, and that even if recursive self-improvement matters, there is no obvious lab milestone that suddenly makes human work disappear. The interesting addition here is not blanket dismissal of RSI, but a push to interrogate the discourse assumptions around it while shifting attention toward how work …</description>
      </item>
    
      <item>
        <title>Experimental evidence of recursive self-improvement</title>
        <link>https://reading-list.oddship.net/notes/2026-07-15-experimental-evidence-of-recursive-self-improvement/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-15-experimental-evidence-of-recursive-self-improvement/</guid>
        <pubDate>Wed, 15 Jul 2026 14:02:00 +0530</pubDate>
        <description>Logged at IST: 2026-07-15 14:02 IST
What it is: Zhengyao Jiang claiming the first experimental evidence of recursive self-improvement in an autoresearch agent
Gist: The specific claim is not generic &amp;amp;quot;agents got better with more tuning,&amp;amp;quot; but that an agent spent eight days autoresearching its own harness and produced a variant that beat a hand-tuned baseline built over two years on held-out benchmarks. If the thread substantiates it, the interesting part is not self-modification in the abstract but search over agent workflows yielding benchmark gains that transfer beyond the optimization loop.
N…</description>
      </item>
    
      <item>
        <title>Goel summarizing Deep SWE 1.1 model-cost comparisons</title>
        <link>https://reading-list.oddship.net/notes/2026-07-11-goel-summarizing-deep-swe-1-1-model-cost-comparisons/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-11-goel-summarizing-deep-swe-1-1-model-cost-comparisons/</guid>
        <pubDate>Sat, 11 Jul 2026 00:00:00 +0000</pubDate>
        <description>Logged at IST: 2026-07-11 01:15 IST
What it is: X post by Shantanu Goel summarizing Deep SWE 1.1 model-cost comparisons
Gist: Claims GPT 5.6 Sol medium outperforms Opus 4.8 max at roughly one-sixth the cost, while GPT 5.6 Sol High performs similarly to Fable 5 max at roughly one-fifth the cost. Framed as a benchmark-driven price&amp;amp;#x2F;performance argument rather than a qualitative workflow review.
Newsletter angle: Useful datapoint for coding-model market structure: if these Deep SWE 1.1 comparisons hold up, the story is not just capability but a sharp shift in price-performance for SWE-oriented mod…</description>
      </item>
    
      <item>
        <title>GPT-5.4 with Pi 0.69.0 is just nice</title>
        <link>https://reading-list.oddship.net/notes/2026-07-11-gpt-5-4-with-pi-0-69-0-is-just-nice/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-11-gpt-5-4-with-pi-0-69-0-is-just-nice/</guid>
        <pubDate>Sat, 11 Jul 2026 00:00:00 +0000</pubDate>
        <description>Logged at IST: 2026-07-11 02:17 IST
What it is: X post by Rohan Verma linking his blog post &amp;amp;quot;GPT-5.4 with Pi 0.69.0 is just nice&amp;amp;quot;
Gist: Argues that an agent harness stack getting boring is a success condition, not a failure. The post frames Pi 0.69.0 + GPT-5.4 + Bosun&amp;amp;#x2F;Zero Agent as having crossed from fragile novelty into dependable daily tooling, where the interesting result is not frontier-model hype but the fact that the stack stopped demanding constant maintenance to remain useful.
Newsletter angle: Strong firsthand writeup on harness maturity: the real milestone is when the agent stack st…</description>
      </item>
    
      <item>
        <title>Harness Engineering for Self-Improvement</title>
        <link>https://reading-list.oddship.net/notes/2026-07-11-harness-engineering-for-self-improvement/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-11-harness-engineering-for-self-improvement/</guid>
        <pubDate>Sat, 11 Jul 2026 00:00:00 +0000</pubDate>
        <description>Logged at IST: 2026-07-11 15:38 IST
What it is: Lilian Weng blog post, &amp;amp;quot;Harness Engineering for Self-Improvement&amp;amp;quot;
Gist: Argues that recursive self-improvement in the near term is less about models rewriting their own weights and more about improving the surrounding harness: workflow loops, context management, filesystem memory, subagents, backend jobs, evaluation, and runtime design. The core claim is that the deployment layer between model and world is becoming an optimization target in its own right.
Newsletter angle: Strong framing for why the interesting frontier is shifting from prompt tr…</description>
      </item>
    
      <item>
        <title>Hashimoto on side-by-side Sol xhigh versus Ultra runs</title>
        <link>https://reading-list.oddship.net/notes/2026-07-11-hashimoto-on-side-by-side-sol-xhigh-versus-ultra-runs/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-11-hashimoto-on-side-by-side-sol-xhigh-versus-ultra-runs/</guid>
        <pubDate>Sat, 11 Jul 2026 00:00:00 +0000</pubDate>
        <description>Logged at IST: 2026-07-11 01:09 IST
What it is: X post by Mitchell Hashimoto on side-by-side Sol xhigh versus Ultra runs
Gist: Says two days of side-by-side planning and implementation runs did not reveal a tangible quality difference between Sol xhigh and Ultra, even though execution behavior and token usage clearly differed. The underlying question is what real use case, if any, currently justifies paying for the more expensive tier.
Newsletter angle: Good practitioner datapoint on frontier-model tiering: users may see visible cost and execution differences before they see reliable quality s…</description>
      </item>
    
      <item>
        <title>Great Divergence in Software Engineering</title>
        <link>https://reading-list.oddship.net/notes/2026-07-10-great-divergence-in-software-engineering/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-10-great-divergence-in-software-engineering/</guid>
        <pubDate>Fri, 10 Jul 2026 00:00:00 +0000</pubDate>
        <description>Logged at IST: 2026-07-10 12:11 IST
What it is: X post by Geoffrey Huntley linking to Stack72&amp;amp;#x27;s essay &amp;amp;quot;The Great Divergence in Software Engineering&amp;amp;quot;
Gist: Argues that the gap between teams effectively using AI and teams still piloting or rejecting it is no longer a simple lead but a compounding divergence, driven by retooling workflows, encoding automation, and treating bad AI output as an engineering problem instead of a veto.
Newsletter angle: Strong framing for AI-native engineering orgs versus incumbents stuck in evaluation loops; good organizational&amp;amp;#x2F;process lens.
Embedded source

  
    X…</description>
      </item>
    
      <item>
        <title>Humans Are Just Stochastic Parrots</title>
        <link>https://reading-list.oddship.net/notes/2026-07-10-humans-are-just-stochastic-parrots/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-10-humans-are-just-stochastic-parrots/</guid>
        <pubDate>Fri, 10 Jul 2026 00:00:00 +0000</pubDate>
        <description>Logged at IST: 2026-07-10 11:07 IST
What it is: Ryan Dahl essay, &amp;amp;quot;Humans Are Just Stochastic Parrots&amp;amp;quot;
Gist: A satirical inversion of common anti-LLM critiques, applying them to humans to highlight how shallow many stochastic-parrot arguments are when stripped of their double standard.
Newsletter angle: Sharp rhetorical piece in the AI discourse wars; useful as culture&amp;amp;#x2F;argumentation rather than technical substance.
</description>
      </item>
    
      <item>
        <title>long talk by the ex-NVIDIA engineer behind Unsloth on fine-tuning and reasoning-model workflows</title>
        <link>https://reading-list.oddship.net/notes/2026-07-10-long-talk-by-the-ex-nvidia-engineer-behind-unsloth-on-fine-tuning-and-reasoning-model-workflows/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-10-long-talk-by-the-ex-nvidia-engineer-behind-unsloth-on-fine-tuning-and-reasoning-model-workflows/</guid>
        <pubDate>Fri, 10 Jul 2026 00:00:00 +0000</pubDate>
        <description>Logged at IST: 2026-07-10 00:05 IST
What it is: X post by h100envy summarizing a long talk by the ex-NVIDIA engineer behind Unsloth on fine-tuning and reasoning-model workflows
Gist: Frames a practical single-GPU stack for local&amp;amp;#x2F;post-training work: choose a base model, use Triton kernels for faster fine-tuning, quantize to 4-bit, run GRPO&amp;amp;#x2F;DPO, and ship a reasoning model on hardware you already own.
Newsletter angle: Useful pointer for the current small team &amp;amp;#x2F; single GPU post-training stack around Unsloth, Triton, quantization, and RLHF-style methods.
Retrieval note: I could ground this from th…</description>
      </item>
    
      <item>
        <title>Cloudflare blog post introducing Meerkat, a new global consensus service built on the QuePaxa algorithm</title>
        <link>https://reading-list.oddship.net/notes/2026-07-09-cloudflare-blog-post-introducing-meerkat-a-new-global-consensus-service-built-on-the-quepaxa-algori/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-09-cloudflare-blog-post-introducing-meerkat-a-new-global-consensus-service-built-on-the-quepaxa-algori/</guid>
        <pubDate>Thu, 09 Jul 2026 00:00:00 +0000</pubDate>
        <description>Logged at IST: 2026-07-09 23:10 IST
What it is: Cloudflare blog post introducing Meerkat, a new global consensus service built on the QuePaxa algorithm
Gist: Cloudflare is building Meerkat for strongly consistent control-plane state across 330+ data centers, arguing that leader-and-timeout-heavy approaches like Raft are a poor fit for hostile WAN conditions and that QuePaxa’s all-replicas-can-write model better matches their network.
Newsletter angle: Notable systems&amp;amp;#x2F;infrastructure piece on consensus design beyond Raft, especially for globally distributed control planes.
</description>
      </item>
    
      <item>
        <title>Some new agentic patterns</title>
        <link>https://reading-list.oddship.net/notes/2026-07-08-some-new-agentic-patterns/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-08-some-new-agentic-patterns/</guid>
        <pubDate>Wed, 08 Jul 2026 00:00:00 +0000</pubDate>
        <description>Logged at IST: 2026-07-08 22:40 IST
What it is: X post by Bilgin Ibryam linking to Prime Radiant&amp;amp;#x27;s &amp;amp;quot;Some new agentic patterns&amp;amp;quot;
Gist: Describes production-ish internal agent patterns built around an &amp;amp;quot;agentic user in the loop&amp;amp;quot; model, with agents in Slack handling intake, ticketing, wiki updates, EA-style assistance, and subagent&amp;amp;#x2F;container-backed workflows.
Newsletter angle: Concrete patterns for embedding agents into team operations without pretending they are fully autonomous replacements.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. C…</description>
      </item>
    
      <item>
        <title>Andrej Jovanović announcing the Red Queen Gödel Machine (arXiv:2606.26294)</title>
        <link>https://reading-list.oddship.net/notes/2026-07-06-andrej-jovanovi-announcing-the-red-queen-g-del-machine-arxiv-2606-26294/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-06-andrej-jovanovi-announcing-the-red-queen-g-del-machine-arxiv-2606-26294/</guid>
        <pubDate>Mon, 06 Jul 2026 00:00:00 +0000</pubDate>
        <description>What it is: Andrej Jovanović announcing the Red Queen Gödel Machine (arXiv:2606.26294)
Gist: self-improving agents should co-evolve with the evaluators that judge them; otherwise stronger agents just learn to exploit stale tests. Paper claims better coding performance with 1.35x–1.72x fewer tokens plus gains in review&amp;amp;#x2F;grading tasks
Newsletter angle: smarter agents need smarter judges; the judge is becoming part of the frontier
Retrieval note: metadata&amp;amp;#x2F;abstract pulled from arXiv; early reproduction repo found at ianyac&amp;amp;#x2F;red-queen-godel-machine
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show em…</description>
      </item>
    
      <item>
        <title>Harness Engineering for Self-Improvement</title>
        <link>https://reading-list.oddship.net/notes/2026-07-06-harness-engineering-for-self-improvement/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-06-harness-engineering-for-self-improvement/</guid>
        <pubDate>Mon, 06 Jul 2026 00:00:00 +0000</pubDate>
        <description>What it is: Lilian Weng sharing her new Lil&amp;amp;#x27;Log post, &amp;amp;quot;Harness Engineering for Self-Improvement&amp;amp;quot;
Gist: argues recursive self-improvement will depend not just on better base models but on better harnesses, the runtime layer that manages tools, planning loops, context, permissions, persistent files, evaluation, and subagents. Strong recurring patterns are workflow automation, file-system-backed persistent memory, and explicit parallel subagent&amp;amp;#x2F;job management
Newsletter angle: the real frontier in RSI may be the software system around the model, not just the model weights themselves
Retrieval not…</description>
      </item>
    
      <item>
        <title>lutke linking to a new Evolution paper by Steven A. Frank</title>
        <link>https://reading-list.oddship.net/notes/2026-07-06-lutke-linking-to-a-new-evolution-paper-by-steven-a-frank/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-06-lutke-linking-to-a-new-evolution-paper-by-steven-a-frank/</guid>
        <pubDate>Mon, 06 Jul 2026 00:00:00 +0000</pubDate>
        <description>What it is: X post by tobi lutke linking to a new Evolution paper by Steven A. Frank
Gist: paper argues evolvability is best understood as generalization; imports modern ML intuition that larger &amp;amp;#x2F; more parameterized systems can generalize better, then maps that onto biological complexity and genomic&amp;amp;#x2F;regulatory capacity
Newsletter angle: evolution-as-generalization; complexity as reusable-solution capacity rather than mere accumulation
Retrieval note: X content recovered via oEmbed; destination paper metadata&amp;amp;#x2F;abstract&amp;amp;#x2F;context reconstructed from Crossref + OpenAlex because publisher page was bot…</description>
      </item>
    
      <item>
        <title>Maxime Rivest demo turning a reMarkable Paper Pro into Tom Riddle’s diary</title>
        <link>https://reading-list.oddship.net/notes/2026-07-06-maxime-rivest-demo-turning-a-remarkable-paper-pro-into-tom-riddle-s-diary/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-06-maxime-rivest-demo-turning-a-remarkable-paper-pro-into-tom-riddle-s-diary/</guid>
        <pubDate>Mon, 06 Jul 2026 00:00:00 +0000</pubDate>
        <description>What it is: Maxime Rivest demo turning a reMarkable Paper Pro into Tom Riddle’s diary
Gist: real hardware&amp;amp;#x2F;software hack where handwritten ink fades, an on-device LLM reads the page, and replies animate back in handwriting; strongest idea is AI as object&amp;amp;#x2F;interface magic rather than another chat box
Newsletter angle: narrative interfaces; AI gets more compelling when wrapped in a strong object metaphor
Retrieval note: X media post grounded in the linked GitHub repo&amp;amp;#x2F;README (MaximeRivest&amp;amp;#x2F;riddle)
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit…</description>
      </item>
    
      <item>
        <title>Nithin Kamath on Zerodha’s operating culture</title>
        <link>https://reading-list.oddship.net/notes/2026-07-06-nithin-kamath-on-zerodha-s-operating-culture/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-06-nithin-kamath-on-zerodha-s-operating-culture/</guid>
        <pubDate>Mon, 06 Jul 2026 00:00:00 +0000</pubDate>
        <description>What it is: Nithin Kamath on Zerodha’s operating culture
Gist: a nice place to work is not accidental culture but the result of repeated leadership choices, slowing down to avoid burnout, staying small, avoiding fear-based management, letting tech make technical decisions, and refusing toxic revenue incentives
Notable line: “A nice place to work is not a perk we offer. It is kind of a business model in itself.”
Newsletter angle: culture as operating system &amp;amp;#x2F; business model, not HR perk
Retrieval note: read directly from source post
</description>
      </item>
    
      <item>
        <title>agent-driven testing and token cost tradeoffs between text buffers and screenshots</title>
        <link>https://reading-list.oddship.net/notes/2026-07-05-agent-driven-testing-and-token-cost-tradeoffs-between-text-buffers-and-screenshots/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-05-agent-driven-testing-and-token-cost-tradeoffs-between-text-buffers-and-screenshots/</guid>
        <pubDate>Sun, 05 Jul 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

X post by @jlongster on agent-driven testing and token cost tradeoffs between text buffers and screenshots.
Gist: for app-testing agents, screenshots are surprisingly close to text buffers on token cost in some models, but vary a lot by provider&amp;amp;#x2F;model; OpenAI looks relatively cheap for images in his comparison while Anthropic is notably higher.
Why it matters: useful for designing UI&amp;amp;#x2F;testing agents without assuming vision is prohibitively expensive.
Newsletter angle: “vision for agentic testing may already be economically viable, depending on model choice…</description>
      </item>
    
      <item>
        <title>replying to @_svs_, quoting Neil Gaiman’s 2013 Guardian piece on libraries&#x2F;reading&#x2F;daydreaming</title>
        <link>https://reading-list.oddship.net/notes/2026-07-05-replying-to-svs-quoting-neil-gaiman-s-2013-guardian-piece-on-libraries-reading-daydreaming/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-05-replying-to-svs-quoting-neil-gaiman-s-2013-guardian-piece-on-libraries-reading-daydreaming/</guid>
        <pubDate>Sun, 05 Jul 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

X post by @captn3m0 replying to @svs, quoting Neil Gaiman’s 2013 Guardian piece on libraries&amp;amp;#x2F;reading&amp;amp;#x2F;daydreaming.
Gist: pushback on using science fiction as an interpretive frame for AI; claim is SF is valuable for dreams&amp;amp;#x2F;imagination, but bad as an answer-book or analogy source for present AI.
Why it matters: clean distinction between fiction as imagination engine vs fiction as policy&amp;amp;#x2F;analysis substrate.
Newsletter angle: “stop using sci-fi as AI governance shorthand” &amp;amp;#x2F; tension between imagination’s value and analogy overreach.
Retrieval note: extracted p…</description>
      </item>
    
      <item>
        <title>Read More (Science) Fiction</title>
        <link>https://reading-list.oddship.net/notes/2026-07-04-read-more-science-fiction/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-04-read-more-science-fiction/</guid>
        <pubDate>Sat, 04 Jul 2026 00:00:00 +0000</pubDate>
        <description>What it is: X post from svs sharing his essay “Read More (Science) Fiction.”
Newsletter angle: “read more sci-fi” is the visible conclusion, but the sharper claim is that fiction supplies vocab and priors for handling agentic weirdness without naive hype or naive panic.
Retrieval note: extracted via FXTwitter API + fetched linked article directly.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once to load X embeds on this site; this choice is remembered in this browser.
  
  Open post on X instead
  Open post on X

</description>
      </item>
    
      <item>
        <title>Should LLMs just treat text content as an image?</title>
        <link>https://reading-list.oddship.net/notes/2026-07-04-should-llms-just-treat-text-content-as-an-image/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-04-should-llms-just-treat-text-content-as-an-image/</guid>
        <pubDate>Sat, 04 Jul 2026 00:00:00 +0000</pubDate>
        <description>What it is: X reply from Michigan TypeScript pointing to Sean Goedecke’s post “Should LLMs just treat text content as an image?”
Newsletter angle: counterintuitive interface hack + deeper architectural question about whether text should sometimes ride the vision path.
Retrieval note: extracted via FXTwitter API + fetched linked article directly.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once to load X embeds on this site; this choice is remembered in this browser.
  
  Open post on X instead
  Open post on X

</description>
      </item>
    
      <item>
        <title>about Passmark, an open-source Playwright library for AI browser regression testing</title>
        <link>https://reading-list.oddship.net/notes/2026-07-01-about-passmark-an-open-source-playwright-library-for-ai-browser-regression-testing/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-01-about-passmark-an-open-source-playwright-library-for-ai-browser-regression-testing/</guid>
        <pubDate>Wed, 01 Jul 2026 00:00:00 +0000</pubDate>
        <description>Gist: Passmark checks an LLM cache before making fresh model calls during browser tests, reportedly cutting a 7-minute suite down to 90 seconds.
Newsletter angle: “cache-first AI browser testing” as practical infra for regression pipelines.
Retrieval note: tweet extracted via FXTwitter API; attached screenshot shows the GitHub repo tagline mentioning intelligent caching, authentication, and multi-model verification.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once to load X embeds on this site; this choice is remembered in this…</description>
      </item>
    
      <item>
        <title>Ponnappa sharing a Realfast blog post by Harsh Jain</title>
        <link>https://reading-list.oddship.net/notes/2026-07-01-ponnappa-sharing-a-realfast-blog-post-by-harsh-jain/</link>
        <guid>https://reading-list.oddship.net/notes/2026-07-01-ponnappa-sharing-a-realfast-blog-post-by-harsh-jain/</guid>
        <pubDate>Wed, 01 Jul 2026 00:00:00 +0000</pubDate>
        <description>Gist: the claim is that in large legacy systems, the bottleneck is not typing code but building enough system understanding to change it safely and quickly; LLMs compress that comprehension step.
Newsletter angle: “AI helps most where system understanding dominates implementation” with a grounded enterprise-delivery example.
Retrieval note: extracted via FXTwitter API; linked article title&amp;amp;#x2F;card also reinforce the same point (“Velocity isn&amp;amp;#x27;t lines per day. It&amp;amp;#x27;s knowing which lines matter.”).
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit.…</description>
      </item>
    
      <item>
        <title>Profiling | Internals for Interns</title>
        <link>https://reading-list.oddship.net/notes/2026-06-30-profiling-internals-for-interns/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-30-profiling-internals-for-interns/</guid>
        <pubDate>Tue, 30 Jun 2026 00:00:00 +0000</pubDate>
        <description>Gist: all five profiles emit the same pprof structure; the core difference is collection model, CPU samples asynchronously via signal + ring buffer, heap&amp;amp;#x2F;block&amp;amp;#x2F;mutex aggregate in per-stack tables in place, goroutine snapshots stacks on demand.
Newsletter angle: “pprof is one file format over three collection strategies” is a clean framing hook.
Retrieval note: extracted via FXTwitter API + linked article fetch.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once to load X embeds on this site; this choice is remembered in this brow…</description>
      </item>
    
      <item>
        <title>Orosz linking Semgrep’s benchmark writeup on GLM 5.2 vs Claude for IDOR detection</title>
        <link>https://reading-list.oddship.net/notes/2026-06-29-orosz-linking-semgrep-s-benchmark-writeup-on-glm-5-2-vs-claude-for-idor-detection/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-29-orosz-linking-semgrep-s-benchmark-writeup-on-glm-5-2-vs-claude-for-idor-detection/</guid>
        <pubDate>Mon, 29 Jun 2026 00:00:00 +0000</pubDate>
        <description>Gist: on Semgrep’s IDOR benchmark, GLM 5.2 scored 39% F1 in a simple prompt-only PydanticAI harness, beating Claude Code’s 32% while costing roughly $0.17 per vulnerability found; Semgrep’s own endpoint-discovery multimodal harness still led overall at 53–61% F1.
Newsletter angle: “the harness matters more than the model, until a cheap open model gets good enough to change the default stack.”
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once to load X embeds on this site; this choice is remembered in this browser.
  
  Open post…</description>
      </item>
    
      <item>
        <title>Orosz quoting Brian Armstrong on Coinbase AI infra economics</title>
        <link>https://reading-list.oddship.net/notes/2026-06-28-orosz-quoting-brian-armstrong-on-coinbase-ai-infra-economics/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-28-orosz-quoting-brian-armstrong-on-coinbase-ai-infra-economics/</guid>
        <pubDate>Sun, 28 Jun 2026 00:00:00 +0000</pubDate>
        <description>Gist: Coinbase reportedly cut AI spend nearly in half while token usage kept growing by changing defaults to cheaper open-weight models (GLM 5.2, Kimi 2.7), adding smarter routing, aggressively using caching, and keeping context lean instead of tightening caps.
Newsletter angle: “AI cost control is becoming a systems problem, not a policy problem.”
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once to load X embeds on this site; this choice is remembered in this browser.
  
  Open post on X instead
  Open post on X

</description>
      </item>
    
      <item>
        <title>slime</title>
        <link>https://reading-list.oddship.net/notes/2026-06-28-slime/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-28-slime/</guid>
        <pubDate>Sun, 28 Jun 2026 00:00:00 +0000</pubDate>
        <description>Gist: the design claim is “one stable RL kernel, task-specific variety in data generation.” Training stays fixed; multi-turn tools, environment feedback, verifier rewards, and other agent behaviors are modeled as rollout&amp;amp;#x2F;data-gen differences rather than separate trainer forks.
Newsletter angle: “Agent RL stacks may converge on a small trusted core plus flexible data-generation layers.”
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once to load X embeds on this site; this choice is remembered in this browser.
  
  Open post on X i…</description>
      </item>
    
      <item>
        <title>Bilgin Ibryam sharing an article on Portkey’s product-engineering org design</title>
        <link>https://reading-list.oddship.net/notes/2026-06-26-bilgin-ibryam-sharing-an-article-on-portkey-s-product-engineering-org-design/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-26-bilgin-ibryam-sharing-an-article-on-portkey-s-product-engineering-org-design/</guid>
        <pubDate>Fri, 26 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Bilgin Ibryam sharing an article on Portkey’s product-engineering org design
Gist: highlights a notably lean product org, 24 product engineers, 1 product designer, 0 PMs, and frames the build&amp;amp;#x2F;operating model as the interesting part
Newsletter angle: “the product engineer company” &amp;amp;#x2F; what gets easier or riskier when PM functions collapse into eng
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once to load X embeds on this site; this choice is remembered in this browser.
  
  Open post on X instead
  Open post on X

</description>
      </item>
    
      <item>
        <title>Go&#x2F;security post on building a self-hosted LLM security proxy with sub-2ms prompt inspection</title>
        <link>https://reading-list.oddship.net/notes/2026-06-26-go-security-post-on-building-a-self-hosted-llm-security-proxy-with-sub-2ms-prompt-inspection/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-26-go-security-post-on-building-a-self-hosted-llm-security-proxy-with-sub-2ms-prompt-inspection/</guid>
        <pubDate>Fri, 26 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Go&amp;amp;#x2F;security post on building a self-hosted LLM security proxy with sub-2ms prompt inspection
Gist: author built an OpenAI-compatible reverse proxy (“Tamga”) that scans prompts for PII, secrets, and prompt-injection patterns before forwarding to providers; key engineering lesson is a hybrid scan pipeline where cheap CPU-bound detectors run sequentially while slower network&amp;amp;#x2F;model-backed scanners run in parallel, because goroutine orchestration overhead dominated when everything fanned out
Newsletter angle: concrete infra pattern for “LLM middleware” that is more about latency budgets…</description>
      </item>
    
      <item>
        <title>How I use LLMs as a staff engineer in 2026</title>
        <link>https://reading-list.oddship.net/notes/2026-06-26-how-i-use-llms-as-a-staff-engineer-in-2026/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-26-how-i-use-llms-as-a-staff-engineer-in-2026/</guid>
        <pubDate>Fri, 26 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Bilgin Ibryam sharing Sean Goedecke’s updated “How I use LLMs as a staff engineer in 2026” workflow writeup
Gist: the notable shift versus 2025 is treating agents as default collaborators for nearly every code change, bug investigation, codebase research, testing, and local setup, while still keeping humans responsible for review, judgment, PR descriptions, ADRs&amp;amp;#x2F;messages, and UI evaluation; especially strong on the idea that current agents are now good enough to generate full PRs and chase bugs across repos, but still need selection, steering, and rejection by an experienced engine…</description>
      </item>
    
      <item>
        <title>Kenton Varda argues against per-agent manual permission configuration and for capability-based security for...</title>
        <link>https://reading-list.oddship.net/notes/2026-06-24-kenton-varda-argues-against-per-agent-manual-permission-configuration-and-for-capability-based-secu/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-24-kenton-varda-argues-against-per-agent-manual-permission-configuration-and-for-capability-based-secu/</guid>
        <pubDate>Wed, 24 Jun 2026 00:00:00 +0000</pubDate>
        <description>Gist: the safe&amp;amp;#x2F;scalable model is many fine-grained task-specific agents, each receiving only the exact capabilities implied by the task context (for example, a pasted doc URL grants access only to that doc). He also argues agent authority should derive from a human principal for accountability, and team-shared setups should be reproducible under each user’s credentials.
Newsletter angle: capability security as the missing abstraction for practical agent authorization; good counterpoint to broad workspace-level agent identity models.
Retrieval note: extracted via FXTwitter API note tweet text; …</description>
      </item>
    
      <item>
        <title>Visible standouts: The Second Half; Eugene Yan on eval process; Han-Chung Lee on agent eval infra; Hamel&#x2F;Sh...</title>
        <link>https://reading-list.oddship.net/notes/2026-06-24-visible-standouts-the-second-half-eugene-yan-on-eval-process-han-chung-lee-on-agent-eval-infra-hame/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-24-visible-standouts-the-second-half-eugene-yan-on-eval-process-han-chung-lee-on-agent-eval-infra-hame/</guid>
        <pubDate>Wed, 24 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Visible standouts: The Second Half; Eugene Yan on eval process; Han-Chung Lee on agent eval infra; Hamel&amp;amp;#x2F;Shreya LLM Evals FAQ; Jason Wei on verification; Anthropic on agent evals; Ofir Press on benchmarks; AI Agents That Matter; Building on Evaluation Quicksand; EvalGen; Benches 2026.
Gist: strong starter pack for agent&amp;amp;#x2F;LLM evals; themes include eval infra as technical debt, process over tooling, verifier design, benchmark saturation&amp;amp;#x2F;contamination, agent-specific eval design, and criteria drift.
Newsletter angle: compact “best evals reading list” &amp;amp;#x2F; why eval practice is shifting fro…</description>
      </item>
    
      <item>
        <title>Coming Loop</title>
        <link>https://reading-list.oddship.net/notes/2026-06-23-coming-loop/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-23-coming-loop/</guid>
        <pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Armin Ronacher post linking to “The Coming Loop”
Gist: argues the important new layer in coding agents is the harness-level loop outside the agent itself; loops already work well for bounded, verifiable work like ports, benchmarking, scanning, and research, but he’s skeptical of using them to write long-lived code because they amplify defensive&amp;amp;#x2F;local reasoning, erode strong invariants, and reduce human comprehension.
Newsletter angle: “The harness is the product” &amp;amp;#x2F; why durable task loops are both inevitable and dangerous.
Note: extracted via FXTwitter API + article fetch; article b…</description>
      </item>
    
      <item>
        <title>David Rosenthal on the AI affordability crisis</title>
        <link>https://reading-list.oddship.net/notes/2026-06-23-david-rosenthal-on-the-ai-affordability-crisis/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-23-david-rosenthal-on-the-ai-affordability-crisis/</guid>
        <pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate>
        <description>Gist: argues model vendors have been massively subsidizing usage to manufacture demand, but token-based pricing is now exposing the real cost structure; for serious enterprise&amp;amp;#x2F;agentic use, compute bills can exceed human labor costs by a wide margin.
Newsletter angle: the agent era may run into a pricing wall before it hits a capability wall.
Note: article extracted successfully via web fetch, though long body was truncated near the footnotes.
</description>
      </item>
    
      <item>
        <title>linking github.com&#x2F;leyten&#x2F;shard</title>
        <link>https://reading-list.oddship.net/notes/2026-06-19-linking-github-com-leyten-shard/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-19-linking-github-com-leyten-shard/</guid>
        <pubDate>Fri, 19 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: X post by @leyten linking github.com&amp;amp;#x2F;leyten&amp;amp;#x2F;shard
Gist: Shard is a WAN-distributed pipeline-parallel LLM inference engine that splits a frontier-size model across GPUs on separate machines; claim is ~30 tok&amp;amp;#x2F;s for GLM-5.2 744B across 6 RTX PRO 6000s in 6 US states using speculative decoding, async pipelining, and a CUDA-graphed draft model.
Newsletter angle: “frontier inference without a datacenter” &amp;amp;#x2F; distributed serving as systems engineering rather than centralized infra.
Notes: extracted via FXTwitter API + GitHub README.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded…</description>
      </item>
    
      <item>
        <title>AI lab business models: subscription vs API</title>
        <link>https://reading-list.oddship.net/notes/2026-06-12-ai-lab-business-models-subscription-vs-api/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-12-ai-lab-business-models-subscription-vs-api/</guid>
        <pubDate>Fri, 12 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: X thread by SemiAnalysis on AI lab business models: subscription vs API.
Gist: based on exhausting weekly limits with real long-horizon coding tasks, they claim consumer subscriptions are far more generous than common “monthly fee ~= API token value ceiling” assumptions. The attached table estimates approximate max monthly value at Claude Pro $20-&amp;amp;amp;gt;$400, Claude Max 5x $100-&amp;amp;amp;gt;$2,000, Claude Max 20x $200-&amp;amp;amp;gt;$8,000; ChatGPT Plus $20-&amp;amp;amp;gt;$700, Pro 5x $100-&amp;amp;amp;gt;$3,500, Pro 20x $200-&amp;amp;amp;gt;$14,000.
Newsletter angle: “AI subscriptions are much more generous than API-pricing intuition sug…</description>
      </item>
    
      <item>
        <title>agent experience</title>
        <link>https://reading-list.oddship.net/notes/2026-06-11-agent-experience/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-11-agent-experience/</guid>
        <pubDate>Thu, 11 Jun 2026 00:00:00 +0000</pubDate>
        <description>Gist: argues DX thinking should extend to agents; optimize the layer between model and codebase via minimal&amp;amp;#x2F;tested context, deterministic environments, proof-heavy verification, structural safety, governance&amp;amp;#x2F;model routing, clean codebase interfaces, and shared preview&amp;amp;#x2F;review loops.
Newsletter angle: “AX as the new DX” + practical checklist for repo&amp;amp;#x2F;runtime&amp;amp;#x2F;review design.
Retrieval note: extracted via FXTwitter API; followed linked Builder article for full gist.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once to load X embeds o…</description>
      </item>
    
      <item>
        <title>Narayanan pointing to a Normal Tech essay on why AI hasn’t replaced software engineers</title>
        <link>https://reading-list.oddship.net/notes/2026-06-11-narayanan-pointing-to-a-normal-tech-essay-on-why-ai-hasn-t-replaced-software-engineers/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-11-narayanan-pointing-to-a-normal-tech-essay-on-why-ai-hasn-t-replaced-software-engineers/</guid>
        <pubDate>Thu, 11 Jun 2026 00:00:00 +0000</pubDate>
        <description>Gist: argues the “AI is replacing software engineers” story is mostly AI-washed layoffs rather than evidence of capability-driven displacement; software work is a decide-execute-deliver sandwich, and AI mainly compresses the execute middle while decision-making, accountability, and deep contextual understanding remain stubbornly human bottlenecks.
Newsletter angle: “AI writes more code, but that’s not the same as replacing engineers” + sandwich model &amp;amp;#x2F; anti-AI-washing thesis.
Retrieval note: extracted via FXTwitter API; followed linked Normal Tech essay for fuller argument (article fetch trunc…</description>
      </item>
    
      <item>
        <title>Code as Agent Harness</title>
        <link>https://reading-list.oddship.net/notes/2026-06-10-code-as-agent-harness/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-10-code-as-agent-harness/</guid>
        <pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: How To AI thread summarizing the Stanford + Meta “Code as Agent Harness” paper.
Gist: the core claim is that reliable agents should externalize reasoning into executable code instead of relying on free-form natural-language chain-of-thought. In this framing, code becomes the agent harness: scripts hold state, tests&amp;amp;#x2F;verifiers provide feedback, execution logs become memory, and the environment constrains behavior through real runtime errors rather than vague self-talk.
Newsletter angle: “the important unit of agent capability is the harness, not the prompt” or “code is becoming the r…</description>
      </item>
    
      <item>
        <title>Deedy post listing standout Claude Fable 5 demos and benchmark anecdotes</title>
        <link>https://reading-list.oddship.net/notes/2026-06-10-deedy-post-listing-standout-claude-fable-5-demos-and-benchmark-anecdotes/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-10-deedy-post-listing-standout-claude-fable-5-demos-and-benchmark-anecdotes/</guid>
        <pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Deedy post listing standout Claude Fable 5 demos and benchmark anecdotes.
Gist: a high-signal hype&amp;amp;#x2F;market snapshot: claims Fable 5 is showing startling capability across large-scale code migration, graphics generation, gameplay, and optimization tasks, while landing near GPT 5.5 pricing. The subtext is that frontier model capability may be moving faster than many software orgs are prepared for.
Newsletter angle: “capability shock is becoming a product-management problem” or “the frontier discourse is shifting from whether to how fast.”
Note: extracted tweet via FXTwitter; no linked…</description>
      </item>
    
      <item>
        <title>Designing loops with Fable 5</title>
        <link>https://reading-list.oddship.net/notes/2026-06-10-designing-loops-with-fable-5/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-10-designing-loops-with-fable-5/</guid>
        <pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: dosco sharing Lance Martin’s “Designing loops with Fable 5”.
Gist: argues stronger agent performance comes from loop design, not just model quality: use explicit goals&amp;amp;#x2F;rubrics for self-correction, separate verifier sub-agents instead of self-critique, and durable memory across sessions. In Lance’s examples, Fable 5 outperformed earlier models by making larger structural bets and benefiting from independent grading plus memory.
Newsletter angle: “better agents need better loops, not just better models” or “independent verification beats self-critique.”
Note: extracted via FXTwitter …</description>
      </item>
    
      <item>
        <title>Eli Bendersky on starting new projects with LLM agents, based on building a new Go project from scratch</title>
        <link>https://reading-list.oddship.net/notes/2026-06-10-eli-bendersky-on-starting-new-projects-with-llm-agents-based-on-building-a-new-go-project-from-scra/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-10-eli-bendersky-on-starting-new-projects-with-llm-agents-based-on-building-a-new-go-project-from-scra/</guid>
        <pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Eli Bendersky on starting new projects with LLM agents, based on building a new Go project from scratch.
Gist: argues agent-heavy development works best when humans keep tight control over design, review, and commit boundaries: start with repo-committed design notes, keep CLs small and reviewable, use strong external tests, and avoid vibe-coding for projects you intend to maintain. He also makes the case that Go is especially agent-friendly because human time shifts from writing to reading.
Newsletter angle: “agent coding turns programming into a reading-heavy discipline” or “small…</description>
      </item>
    
      <item>
        <title>skepticism: the thread oversold it a bit: the paper is a broad survey&#x2F;position piece, not a clean proof th...</title>
        <link>https://reading-list.oddship.net/notes/2026-06-10-skepticism-the-thread-oversold-it-a-bit-the-paper-is-a-broad-survey-position-piece-not-a-clean-proo/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-10-skepticism-the-thread-oversold-it-a-bit-the-paper-is-a-broad-survey-position-piece-not-a-clean-proo/</guid>
        <pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: skepticism: the thread oversold it a bit, the paper is a broad survey&amp;amp;#x2F;position piece, not a clean proof that one architecture flips everything.
Gist: this is mostly a taxonomy and research agenda, not a new experimental result. The paper’s useful move is to separate three layers: code as interface (reasoning, acting, environment modeling), code-enabled harness mechanisms (planning, memory, tool use, plan-execute-verify control, harness optimization), and code as shared substrate for multi-agent coordination. The strongest practical point is that agent reliability lives in the runti…</description>
      </item>
    
      <item>
        <title>agent slop</title>
        <link>https://reading-list.oddship.net/notes/2026-06-09-agent-slop/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-09-agent-slop/</guid>
        <pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Langfuse post&amp;amp;#x2F;article on automating the AI engineering loop without producing “agent slop”.
Gist: argues the whole AI engineering loop can now technically be automated, instrumentation, monitoring, dataset building, testing, deployment, but full automation is a trap when human judgment is the product. Keep humans close to trace review, target definition, and quality-bar decisions.
Newsletter angle: “automate the loop, but not your taste” &amp;amp;#x2F; “agent slop is what happens when evals become the whole target.”
Note: extracted via FXTwitter article payload; fetch was truncated but core arg…</description>
      </item>
    
      <item>
        <title>Decline of Search Engines is an Opportunity</title>
        <link>https://reading-list.oddship.net/notes/2026-06-09-decline-of-search-engines-is-an-opportunity/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-09-decline-of-search-engines-is-an-opportunity/</guid>
        <pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Lewis Campbell post linking to “The Decline of Search Engines is an Opportunity”.
Gist: argues worsening search quality should push people back toward the old web habit of maintaining personal links pages; discovery by human-curated hyperlinks is framed as a healthier alternative to SEO sludge and LLM-mediated search summaries.
Newsletter angle: “search decay revives the links page” is a clean thesis with nice historical texture.
Note: extracted tweet via FXTwitter and fetched linked blog post successfully.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embed…</description>
      </item>
    
      <item>
        <title>Dynamo and the Computer</title>
        <link>https://reading-list.oddship.net/notes/2026-06-09-dynamo-and-the-computer/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-09-dynamo-and-the-computer/</guid>
        <pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Zara Zhang post using Paul David’s “The Dynamo and the Computer” as an analogy for AI adoption.
Gist: argues AI gains won’t come from simply inserting models into existing workflows; like electrification, the real productivity jump comes only after redesigning the organization and flow of work around the new technology.
Newsletter angle: “AI is still in the faster steam engine phase” is a strong line for transformation skepticism.
Note: extracted tweet via FXTwitter; referenced paper link appears to be in replies&amp;amp;#x2F;comments and was not followed here.
Embedded source

  
    X &amp;amp;#x2F; Twitt…</description>
      </item>
    
      <item>
        <title>Our fears about AI are really fears about capitalism</title>
        <link>https://reading-list.oddship.net/notes/2026-06-09-our-fears-about-ai-are-really-fears-about-capitalism/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-09-our-fears-about-ai-are-really-fears-about-capitalism/</guid>
        <pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: LTSE post linking Eric Ries’s Fast Company essay, “Our fears about AI are really fears about capitalism”.
Gist: argues many AI anxieties are really about institutions and incentive systems optimizing for the wrong outcomes; the key question is not just what machines optimize for, but what organizations optimize for.
Newsletter angle: “AI fear is often misdirected systems fear” or “alignment problems are organizational too, not just model-level.”
Note: extracted tweet via FXTwitter; direct article fetch was blocked by Fast Company anti-bot checks, so gist is based on the linked titl…</description>
      </item>
    
      <item>
        <title>Shriram Krishnamurthi memo on rebooting a programming languages course for the agentic coding era</title>
        <link>https://reading-list.oddship.net/notes/2026-06-09-shriram-krishnamurthi-memo-on-rebooting-a-programming-languages-course-for-the-agentic-coding-era/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-09-shriram-krishnamurthi-memo-on-rebooting-a-programming-languages-course-for-the-agentic-coding-era/</guid>
        <pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Shriram Krishnamurthi memo on rebooting a programming languages course for the agentic coding era.
Gist: argues PL should be reframed around constraining AI-generated implementations and providing guarantees; distinguishes PL from SE&amp;amp;#x2F;FM, then proposes teaching along two axes: language confinement and custom program properties.
Newsletter angle: “AI makes PL more about guarantees than syntax” + course design as a forecast of curriculum shifts.
Note: extracted tweet via FXTwitter, then fetched linked public Google Doc; captured substantive sections including motivation and course str…</description>
      </item>
    
      <item>
        <title>What is an agent?</title>
        <link>https://reading-list.oddship.net/notes/2026-06-09-what-is-an-agent/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-09-what-is-an-agent/</guid>
        <pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Karthik S sharing Hadley Wickham’s “What is an agent?” explainer.
Gist: very clear mental model: an agent is an LLM inside a harness that can call tools repeatedly in a loop; the harness mediates tool calls&amp;amp;#x2F;results and turns a stateless request&amp;amp;#x2F;response model into iterative action.
Newsletter angle: “agent = looped tool use inside a harness” is a concise definitional anchor for broader agent discussions.
Note: extracted tweet via FXTwitter and fetched linked Substack article successfully.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track y…</description>
      </item>
    
      <item>
        <title>Armin Ronacher explaining Pi’s new per-project approval prompt and the security model behind it</title>
        <link>https://reading-list.oddship.net/notes/2026-06-08-armin-ronacher-explaining-pi-s-new-per-project-approval-prompt-and-the-security-model-behind-it/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-08-armin-ronacher-explaining-pi-s-new-per-project-approval-prompt-and-the-security-model-behind-it/</guid>
        <pubDate>Mon, 08 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Armin Ronacher explaining Pi’s new per-project approval prompt and the security model behind it.
Gist: the key argument is that AGENTS.md gets injected into the system prompt, so untrusted repo-level instructions can directly influence agent behavior in ways a README usually won’t; Pi added one-time trust prompts to reduce silent execution risk on untrusted repos.
Newsletter angle: repo-local agent instructions are becoming both a productivity primitive and a new software supply-chain&amp;amp;#x2F;security surface.
Note: extracted via FXTwitter API from the tweet’s article body; points to GitHu…</description>
      </item>
    
      <item>
        <title>blog essay arguing AI disruption is structurally more threatening to software than many other fields becaus...</title>
        <link>https://reading-list.oddship.net/notes/2026-06-08-blog-essay-arguing-ai-disruption-is-structurally-more-threatening-to-software-than-many-other-field/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-08-blog-essay-arguing-ai-disruption-is-structurally-more-threatening-to-software-than-many-other-field/</guid>
        <pubDate>Mon, 08 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: a blog essay arguing AI disruption is structurally more threatening to software than many other fields because code is verifiable, open source created a huge training corpus, and AI labs can dogfood coding tools on themselves.
Gist: the author expects a race to the bottom in software pricing, wage compression, permanent erosion of the talent pipeline, higher output expectations for remaining engineers, and most upside captured by owners&amp;amp;#x2F;model providers.
Newsletter angle: software may be the cleanest early target for AI because it has both an oracle and the richest public corpus.
No…</description>
      </item>
    
      <item>
        <title>just use loops</title>
        <link>https://reading-list.oddship.net/notes/2026-06-08-just-use-loops/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-08-just-use-loops/</guid>
        <pubDate>Mon, 08 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Gergely Orosz pushing back on the blanket “just use loops” advice for coding agents.
Gist: his claim is that autonomous loop-heavy agent workflows mainly make sense for the relatively small set of people with effectively unlimited token budgets and enough friction with prompt-driven workflows to justify the spend.
Newsletter angle: the real constraint on agent autonomy may be economics, not just capability.
Note: extracted via FXTwitter API from the tweet text only.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once t…</description>
      </item>
    
      <item>
        <title>promoting an 85-minute MIT lecture on Git internals &#x2F; data model</title>
        <link>https://reading-list.oddship.net/notes/2026-06-08-promoting-an-85-minute-mit-lecture-on-git-internals-data-model/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-08-promoting-an-85-minute-mit-lecture-on-git-internals-data-model/</guid>
        <pubDate>Mon, 08 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: an X post promoting an 85-minute MIT lecture on Git internals &amp;amp;#x2F; data model.
Gist: the pitch is that most developers memorize Git commands without understanding commits, trees, refs, and the graph underneath; learning the model makes debugging history and merge&amp;amp;#x2F;rebase failures much less magical.
Newsletter angle: “Git literacy as leverage”, understanding the object graph matters more when agents are branching&amp;amp;#x2F;rewriting history at speed.
Note: extracted via FXTwitter API; saved from the post text only, lecture content itself not yet reviewed.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
…</description>
      </item>
    
      <item>
        <title>Sebastian Raschka summarizing a paper on whether repository-level context files like AGENTS.md actually hel...</title>
        <link>https://reading-list.oddship.net/notes/2026-06-08-sebastian-raschka-summarizing-a-paper-on-whether-repository-level-context-files-like-agents-md-actu/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-08-sebastian-raschka-summarizing-a-paper-on-whether-repository-level-context-files-like-agents-md-actu/</guid>
        <pubDate>Mon, 08 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Sebastian Raschka summarizing a paper on whether repository-level context files like AGENTS.md actually help coding agents.
Gist: in the reported benchmarks, LLM-generated context files were neutral-to-slightly-worse versus no context file, developer-written ones were better than LLM-written ones, and surprisingly the no-context condition was often cheaper&amp;amp;#x2F;more efficient.
Newsletter angle: more agent context is not automatically better, extra instructions can increase exploration cost without improving task success.
Note: extracted from the FXTwitter API article body; links to arXi…</description>
      </item>
    
      <item>
        <title>antirez reacting sharply to Anthropic’s Opus 4.8 as a product&#x2F;management failure rather than a raw model-ca...</title>
        <link>https://reading-list.oddship.net/notes/2026-06-04-antirez-reacting-sharply-to-anthropic-s-opus-4-8-as-a-product-management-failure-rather-than-a-raw/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-04-antirez-reacting-sharply-to-anthropic-s-opus-4-8-as-a-product-management-failure-rather-than-a-raw/</guid>
        <pubDate>Thu, 04 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: antirez reacting sharply to Anthropic’s Opus 4.8 as a product&amp;amp;#x2F;management failure rather than a raw model-capability issue.
Gist: the claim is that shipping a bad model experience is more revealing about product judgment and internal decision-making than about frontier-model feasibility; if quality was not there, not shipping would have been the better move.
Newsletter angle: frontier AI competition may increasingly hinge on release quality and organizational judgment, not just the ceiling of the underlying model.
Note: extracted via FXTwitter API; standalone opinion tweet, no linke…</description>
      </item>
    
      <item>
        <title>Modern Engineering Values,</title>
        <link>https://reading-list.oddship.net/notes/2026-06-04-modern-engineering-values/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-04-modern-engineering-values/</guid>
        <pubDate>Thu, 04 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Christoph Nakazawa sharing his essay “Modern Engineering Values,” framed around Codex as a step-change in developer velocity.
Gist: the piece argues coding is no longer the main bottleneck; the durable values now are strong ownership, taste, strict guardrails with fast feedback loops, repo-local context, stack ownership, and preserving option value while agents do more implementation work.
Newsletter angle: AI doesn’t replace engineering values, it increases the premium on ownership, taste, fast verification, and keeping context where agents can actually use it.
Note: extracted via…</description>
      </item>
    
      <item>
        <title>Solo Climb</title>
        <link>https://reading-list.oddship.net/notes/2026-06-04-solo-climb/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-04-solo-climb/</guid>
        <pubDate>Thu, 04 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Ajey Gore linking his essay “The Solo Climb.”
Gist: the argument is that AI-enabled solo builders and tiny teams only work when they first build a genuinely load-bearing “harness”, trusted tests, evals, specs, and hard gates that can answer “is this safe enough to ship?” without relying on redundant humans.
Newsletter angle: “100x teams” are mostly a harness story, AI leverage scales only when trust, eval, and rollback systems become the new team structure.
Note: extracted via FXTwitter API and linked article; article read partially via web fetch due to truncation, but core thesis …</description>
      </item>
    
      <item>
        <title>Why AI Agents Fail in Production</title>
        <link>https://reading-list.oddship.net/notes/2026-06-04-why-ai-agents-fail-in-production/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-04-why-ai-agents-fail-in-production/</guid>
        <pubDate>Thu, 04 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Bilgin Ibryam pointing to Jani Janakiram’s Diagrid essay “Why AI Agents Fail in Production.”
Gist: the core claim is that agent projects fail less because models are weak and more because teams ship behavior without the production substrate underneath it, especially durability, security&amp;amp;#x2F;identity, cost controls, and observability.
Newsletter angle: the production gap for agents looks a lot like the early microservices gap, the winning layer may be the platform that makes agent workflows restartable, attributable, observable, and cost-bounded.
Note: extracted via FXTwitter API; direc…</description>
      </item>
    
      <item>
        <title>Modern Engineering Values</title>
        <link>https://reading-list.oddship.net/notes/2026-06-03-modern-engineering-values/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-03-modern-engineering-values/</guid>
        <pubDate>Wed, 03 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Christoph Nakazawa’s post on “Modern Engineering Values” and his current LLM-heavy workflow.
Gist: core claims are that coding is no longer the bottleneck, strong guardrails plus tight feedback loops matter more than ever, repo-local context becomes the real operating manual for agents, and small teams with strong ownership&amp;amp;#x2F;taste will outperform larger coordination-heavy orgs.
Newsletter angle: engineering values are being redefined around ownership, taste, guardrails, and context placement rather than raw coding throughput.
Note: extracted via FXTwitter API and linked post; articl…</description>
      </item>
    
      <item>
        <title>Han Xiao on Dataroom, a local-first deep research harness</title>
        <link>https://reading-list.oddship.net/notes/2026-06-02-han-xiao-on-dataroom-a-local-first-deep-research-harness/</link>
        <guid>https://reading-list.oddship.net/notes/2026-06-02-han-xiao-on-dataroom-a-local-first-deep-research-harness/</guid>
        <pubDate>Tue, 02 Jun 2026 00:00:00 +0000</pubDate>
        <description>What it is: Han Xiao on Dataroom, a local-first deep research harness.
Gist: argues deep research should be a cheap, long-running first step for long-horizon tasks; Dataroom uses a small local model on your own GPU, keeps gathering until the package is genuinely comprehensive, and outputs a zip instead of burning frontier-model budget.
Newsletter angle: “local-first deep research” &amp;amp;#x2F; small models + harness design beating expensive frontier calls for the reconnaissance phase.
Note: extracted via FXTwitter API.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X…</description>
      </item>
    
      <item>
        <title>solution might be cancelling my AI subscription</title>
        <link>https://reading-list.oddship.net/notes/2026-05-31-solution-might-be-cancelling-my-ai-subscription/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-31-solution-might-be-cancelling-my-ai-subscription/</guid>
        <pubDate>Sun, 31 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted main post via api.fxtwitter.com fallback, then read linked post directly: https:&amp;amp;#x2F;&amp;amp;#x2F;thoughts.hmmz.org&amp;amp;#x2F;2026-05-31.html
Mario Zechner recommends David&amp;amp;#x27;s post the solution might be cancelling my AI subscription.
Gist: a sharp anti-friction argument against current AI-tool usage patterns, cheap output and minimal resistance can explode side projects, context switching, and pseudo-productivity while degrading attention and commitment.
Why it matters: good counterweight to &amp;amp;quot;more agent throughput = better work&amp;amp;quot; narratives; frames AI as an attention-manag…</description>
      </item>
    
      <item>
        <title>antirez on alternatives to the standard EDIT tool for LLM agents; links to a short blog note</title>
        <link>https://reading-list.oddship.net/notes/2026-05-19-antirez-on-alternatives-to-the-standard-edit-tool-for-llm-agents-links-to-a-short-blog-note/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-19-antirez-on-alternatives-to-the-standard-edit-tool-for-llm-agents-links-to-a-short-blog-note/</guid>
        <pubDate>Tue, 19 May 2026 00:00:00 +0000</pubDate>
        <description>Gist: proposes CAS-style edits using line-number + short checksum tags instead of resending old text verbatim, aiming to save tokens while still guarding against stale or hallucinated edits.
Newsletter angle: &amp;amp;quot;a lighter-weight edit primitive for coding agents: line tags vs full old-text CAS&amp;amp;quot;.
Note: extracted via FXTwitter API + antirez.com post.
Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track your visit. Click once to load X embeds on this site; this choice is remembered in this browser.
  
  Open post on X instead
  Open post on X

</description>
      </item>
    
      <item>
        <title>Project Glasswing: what Mythos showed us</title>
        <link>https://reading-list.oddship.net/notes/2026-05-19-project-glasswing-what-mythos-showed-us/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-19-project-glasswing-what-mythos-showed-us/</guid>
        <pubDate>Tue, 19 May 2026 00:00:00 +0000</pubDate>
        <description>What it is: Cloudflare on testing Anthropic Mythos against 50+ internal repos; links to &amp;amp;quot;Project Glasswing: what Mythos showed us&amp;amp;quot;.
Gist: key claim is that stronger offensive-security models change vuln research from bug spotting to exploit-chain construction and proof generation, but the real bottleneck becomes harness design, triage noise, and scoped parallel workflows rather than just faster patching.
Newsletter angle: &amp;amp;quot;offensive AI doesn&amp;amp;#x27;t just speed up vuln discovery, it forces a redesign of the architecture around triage, coverage, and exploit validation&amp;amp;quot;.
Note: extracted via FXTwitter A…</description>
      </item>
    
      <item>
        <title>Ambitious OSS project pitching WiFi CSI as a privacy-preserving sensing stack: presence detection, breathin...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-13-ambitious-oss-project-pitching-wifi-csi-as-a-privacy-preserving-sensing-stack-presence-detection-br/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-13-ambitious-oss-project-pitching-wifi-csi-as-a-privacy-preserving-sensing-stack-presence-detection-br/</guid>
        <pubDate>Wed, 13 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Ambitious OSS project pitching WiFi CSI as a privacy-preserving sensing stack: presence detection, breathing&amp;amp;#x2F;heart-rate monitoring, activity recognition, rough pose estimation, and through-wall&amp;amp;#x2F;environment sensing using ESP32-S3 nodes.
Interesting angle is the packaging: not just a research demo, but a full “edge intelligence” story with cheap hardware, local processing, attestations, mesh sensing, demos, and a long README translating RF sensing into product language.
Newsletter angle: strong hook if framed as “camera-free spatial intelligence from commod…</description>
      </item>
    
      <item>
        <title>Andras Bacsai jokes that Coolify created a fake repo with fake bounties so agent&#x2F;bot-driven fake PR submiss...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-13-andras-bacsai-jokes-that-coolify-created-a-fake-repo-with-fake-bounties-so-agent-bot-driven-fake-pr/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-13-andras-bacsai-jokes-that-coolify-created-a-fake-repo-with-fake-bounties-so-agent-bot-driven-fake-pr/</guid>
        <pubDate>Wed, 13 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Andras Bacsai jokes that Coolify created a fake repo with fake bounties so agent&amp;amp;#x2F;bot-driven fake PR submissions would self-identify and could be banned from the main repo.
Useful as a sharp anecdote about the emerging spam&amp;amp;#x2F;credibility problem around bounty-chasing coding agents: once PR generation gets cheap, maintainers start building honeypots and authenticity filters.
Newsletter angle: strong, funny hook for a piece on anti-spam countermeasures in the age of agentic OSS contribution.
Retrieval note: extracted via api.fxtwitter.com; gist comes from the …</description>
      </item>
    
      <item>
        <title>Course&#x2F;site on harness engineering for AI coding agents, synthesizing OpenAI + Anthropic guidance into lect...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-13-course-site-on-harness-engineering-for-ai-coding-agents-synthesizing-openai-anthropic-guidance-into/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-13-course-site-on-harness-engineering-for-ai-coding-agents-synthesizing-openai-anthropic-guidance-into/</guid>
        <pubDate>Wed, 13 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Course&amp;amp;#x2F;site on harness engineering for AI coding agents, synthesizing OpenAI + Anthropic guidance into lectures, projects, and ready-to-copy templates.
Core pitch: reliability comes less from a smarter model and more from a closed-loop system, explicit constraints, state management, verification, observability, and control.
Newsletter angle: a useful “meta” resource for the current wave of coding-agent practice, especially good if framing the shift from promptcraft to environment&amp;amp;#x2F;harness design.
Retrieval note: extracted cleanly via web_fetch from the lan…</description>
      </item>
    
      <item>
        <title>Jake’s launch post for sqlite3-parser-js: claims a pure-JS port of SQLite’s parser beats every other JS SQL...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-12-jake-s-launch-post-for-sqlite3-parser-js-claims-a-pure-js-port-of-sqlite-s-parser-beats-every-other/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-12-jake-s-launch-post-for-sqlite3-parser-js-claims-a-pure-js-port-of-sqlite-s-parser-beats-every-other/</guid>
        <pubDate>Tue, 12 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Jake’s launch post for sqlite3-parser-js: claims a pure-JS port of SQLite’s parser beats every other JS SQL parser he benchmarked, including wasm-based options.
Useful companion to the repo itself because the punchline is performance positioning: 2.5x over liteparser, 6x over sqlparser-ts, 10x over node-sql-parser, and much larger gaps vs older parsers.
Newsletter angle: strong “unexpected performance result” framing, pure JS beating wasm competitors for SQL parsing is a nice hook into why parser architecture and generated code shape matter.
Retrieval not…</description>
      </item>
    
      <item>
        <title>JS SQLite parser ported from SQLite’s own Lemon&#x2F;LALR grammar, aimed at being fast, lightweight, browser-fri...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-12-js-sqlite-parser-ported-from-sqlite-s-own-lemon-lalr-grammar-aimed-at-being-fast-lightweight-browse/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-12-js-sqlite-parser-ported-from-sqlite-s-own-lemon-lalr-grammar-aimed-at-being-fast-lightweight-browse/</guid>
        <pubDate>Tue, 12 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

JS SQLite parser ported from SQLite’s own Lemon&amp;amp;#x2F;LALR grammar, aimed at being fast, lightweight, browser-friendly, and more faithful than typical JS SQL parsers.
Notable angle: improved structured diagnostics and hints, plus AST traversal&amp;amp;#x2F;CLI tooling, which makes it more useful for editor tooling, linting, query analysis, or SQL-aware product features.
Newsletter angle: a good example of “serious infra-grade parsing” moving into pure TypeScript without wasm, with a tight value prop around correctness + developer ergonomics.
Retrieval note: extracted from G…</description>
      </item>
    
      <item>
        <title>llm</title>
        <link>https://reading-list.oddship.net/notes/2026-05-12-llm/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-12-llm/</guid>
        <pubDate>Tue, 12 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted via api.fxtwitter.com fallback, then read linked TIL directly: https:&amp;amp;#x2F;&amp;amp;#x2F;til.simonwillison.net&amp;amp;#x2F;llms&amp;amp;#x2F;llm-shebang
Simon Willison shows a neat pattern for using his llm CLI in a shebang line, turning plain-English files or YAML templates into executable scripts.
The more interesting part is not the toy prompt examples but the tool-enabled&amp;amp;#x2F;template-enabled scripts: parameterized prompts, embedded functions, and lightweight agentic shells around LLM&amp;amp;#x2F;tool workflows.
Newsletter angle: a crisp example of LLMs collapsing the boundary between prompt, script…</description>
      </item>
    
      <item>
        <title>translation layer</title>
        <link>https://reading-list.oddship.net/notes/2026-05-12-translation-layer/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-12-translation-layer/</guid>
        <pubDate>Tue, 12 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Blog essay arguing AI compresses the org’s “translation layer” more than any single job title: spec→ticket→PR→release-note work gets cheap, while judgement around why&amp;amp;#x2F;what&amp;amp;#x2F;trust systems gets more valuable.
Strong claim: middle-management and coordination-heavy roles shrink unless they actively contribute to product definition, architecture, evals, or verification.
Newsletter angle: useful framing for how AI changes org shape, not “AI replaces engineers” but “AI eats translation work,” which shifts value toward taste, harnesses, and hands-on decision-maker…</description>
      </item>
    
      <item>
        <title>Saved media locally</title>
        <link>https://reading-list.oddship.net/notes/2026-05-11-saved-media-locally/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-11-saved-media-locally/</guid>
        <pubDate>Mon, 11 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted via api.fxtwitter.com fallback; includes an image illustrating progressive rendering from noise to a clear cat image.
Saved media locally:
Dax reframes coding-agent usage: not like 3D printing one committed layer at a time, but like progressive rendering, start with a blurry whole, then make repeated full passes that sharpen the entire shape.
Follow-up reply worth keeping with it: https:&amp;amp;#x2F;&amp;amp;#x2F;x.com&amp;amp;#x2F;thdxr&amp;amp;#x2F;status&amp;amp;#x2F;2053566249351754193, he says this is actually counter to how his brain naturally imagines construction, which makes the metaphor more intere…</description>
      </item>
    
      <item>
        <title>agent principal-agent problem</title>
        <link>https://reading-list.oddship.net/notes/2026-05-07-agent-principal-agent-problem/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-07-agent-principal-agent-problem/</guid>
        <pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Read The agent principal-agent problem by David Crawshaw.
Core claim: classic review-before-commit code review assumed a human contributor whose effort and understanding could be inferred from the code; agent-mediated contribution breaks that signal and creates a principal-agent problem where reviewers absorb heavy load from low-effort, lightly-validated slop PRs.
The useful distinction is not just agents good&amp;amp;#x2F;bad, but high-trust small teams versus low-trust large organizations: small teams can collapse review and let the human prompter own deployment, wh…</description>
      </item>
    
      <item>
        <title>AI slop</title>
        <link>https://reading-list.oddship.net/notes/2026-05-07-ai-slop/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-07-ai-slop/</guid>
        <pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted main post via api.fxtwitter.com fallback.
Mitchell Hashimoto argues that AI slop is useful as an internal experimentation tool: low-quality generated code&amp;amp;#x2F;UI&amp;amp;#x2F;plugins can dramatically reduce the cost of parallel exploration and API iteration, especially when regeneration is cheaper than careful hand maintenance.
His concrete examples are good: shipping an intentionally rough alpha frontend to focus on core internals, and using overnight agent loops to generate many disposable plugins so the whole ecosystem can be tested before the SDK is stable.
…</description>
      </item>
    
      <item>
        <title>Anthropic&#x27;s core idea is to train a model to verbalize its own internal activations into human-readable tex...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-07-anthropic-s-core-idea-is-to-train-a-model-to-verbalize-its-own-internal-activations-into-human-read/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-07-anthropic-s-core-idea-is-to-train-a-model-to-verbalize-its-own-internal-activations-into-human-read/</guid>
        <pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted the Anthropic post via api.fxtwitter.com fallback and checked the linked research page Natural Language Autoencoders: Turning Claude’s thoughts into text.
Anthropic&amp;amp;#x27;s core idea is to train a model to verbalize its own internal activations into human-readable text, then train a second component to reconstruct the original activation from that explanation; better reconstruction is used as the training signal for better explanations.
This is interesting because it tries to turn interpretability outputs into something directly legible, instead of on…</description>
      </item>
    
      <item>
        <title>Autodata</title>
        <link>https://reading-list.oddship.net/notes/2026-05-07-autodata/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-07-autodata/</guid>
        <pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted main post via api.fxtwitter.com fallback and checked the linked Meta RAM Autodata post plus the referenced justrach&amp;amp;#x2F;devswarm repo and sample issue.
Rach connects her agent workflow to Meta&amp;amp;#x27;s Autodata framing: agents act like data scientists by iterating on a hypothesis, generating data, testing it, validating results, extracting learnings, and then closing the loop.
The linked paper&amp;amp;#x2F;blog&amp;amp;#x27;s core idea is strong: convert inference-time compute into better training&amp;amp;#x2F;eval data quality by having an agent iteratively create data, analyze failures, refin…</description>
      </item>
    
      <item>
        <title>Colossus 1</title>
        <link>https://reading-list.oddship.net/notes/2026-05-07-colossus-1/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-07-colossus-1/</guid>
        <pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted Simon Willison&amp;amp;#x27;s post via api.fxtwitter.com fallback and checked the linked note Notes on the xAI&amp;amp;#x2F;Anthropic data center deal.
Simon&amp;amp;#x27;s main clarification is that Anthropic is getting Colossus 1, while xAI keeps using the larger Colossus 2; early chatter that xAI had given up its own compute was wrong.
The sharper points are around externalities and dependency risk: Colossus 1 reportedly has a particularly bad environmental record, and the deal effectively makes Anthropic dependent on infrastructure controlled by Elon&amp;amp;#x2F;xAI with an explicit we reser…</description>
      </item>
    
      <item>
        <title>DFlash</title>
        <link>https://reading-list.oddship.net/notes/2026-05-07-dflash/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-07-dflash/</guid>
        <pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted main post via api.fxtwitter.com fallback and checked the linked repo z-lab&amp;amp;#x2F;dflash.
Zhijian Liu pitches DFlash for Gemma 4 as an open-source speculative decoding path that can push native Gemma 4 MTP further, claiming up to 6x faster generation at the same quality.
Repo framing: DFlash: Block Diffusion for Flash Speculative Decoding, a lightweight block-diffusion draft model for speculative decoding, with support across Gemma, Qwen, Llama, GPT-OSS, MLX, vLLM, SGLang, and Transformers backends.
What seems notable is not just the speed claim, but t…</description>
      </item>
    
      <item>
        <title>Entire&#x27;s core claim is useful: from ~202k real tool calls across ~1,983 public coding-agent checkpoints, ab...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-07-entire-s-core-claim-is-useful-from-202k-real-tool-calls-across-1-983-public-coding-agent-checkpoint/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-07-entire-s-core-claim-is-useful-from-202k-real-tool-calls-across-1-983-public-coding-agent-checkpoint/</guid>
        <pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted Mario Zechner&amp;amp;#x27;s quote-post via api.fxtwitter.com fallback and checked the linked Entire blog post on agentic search.
Entire&amp;amp;#x27;s core claim is useful: from ~202k real tool calls across ~1,983 public coding-agent checkpoints, about 48.8% were search-related, so search is a first-order agent behavior rather than a side utility.
Their more interesting finding is that raw speed is not the main bottleneck. Making search dramatically faster (ripgrep → fff) only modestly improved end-to-end run time because tool latency was a tiny fraction of total wall c…</description>
      </item>
    
      <item>
        <title>microwavegang</title>
        <link>https://reading-list.oddship.net/notes/2026-05-07-microwavegang/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-07-microwavegang/</guid>
        <pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted main post via api.fxtwitter.com fallback; the quoted tweet and attached screenshot provide the actual context.
Claim: a GPT-3 training loss spike was traced to scraped data from a microwavegang subreddit&amp;amp;#x2F;community full of text like MMMMMMMMMMMMMM and BEEP BEEP BEEP, and the spike disappeared after dataset cleanup.
The screenshot is funny but the underlying lesson is serious: weird narrow-distribution junk data can create visible optimization pathologies, and simple data cleaning can remove dramatic training instability.
Why it matters: this is a…</description>
      </item>
    
      <item>
        <title>Ben Holmes says switching from TipTap&#x2F;ProseMirror to Slate made a rich-text bulleted-list interaction drama...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-06-ben-holmes-says-switching-from-tiptap-prosemirror-to-slate-made-a-rich-text-bulleted-list-interacti/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-06-ben-holmes-says-switching-from-tiptap-prosemirror-to-slate-made-a-rich-text-bulleted-list-interacti/</guid>
        <pubDate>Wed, 06 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted main post via api.fxtwitter.com fallback.
Ben Holmes says switching from TipTap&amp;amp;#x2F;ProseMirror to Slate made a rich-text bulleted-list interaction dramatically easier to build; something that took weeks to half-work in TipTap took a couple of hours in Slate, with Codex helping.
Core signal is less Slate is universally better and more framework ergonomics matter a lot for AI-assisted development: some abstractions are much easier to extend&amp;amp;#x2F;debug with model help.
Why it matters: useful anecdote for editor-stack choice, especially when complex WYSIWYG…</description>
      </item>
    
      <item>
        <title>de</title>
        <link>https://reading-list.oddship.net/notes/2026-05-06-de/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-06-de/</guid>
        <pubDate>Wed, 06 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted main post via api.fxtwitter.com fallback; inspected attached screenshot separately.
Thomas Ptacek argues the .de incident is decisive evidence against DNSSEC as core Internet security functionality.
Attached screenshot captures Cloudflare status text saying it temporarily disabled DNSSEC validation on 1.1.1.1 so .de names would continue resolving while DENIC fixed a DNSSEC signing problem.
Why it matters: this is the sharper, event-driven version of the previous anti-DNSSEC thesis, if a major resolver bypasses validation during a registry signin…</description>
      </item>
    
      <item>
        <title>does not approximate attention</title>
        <link>https://reading-list.oddship.net/notes/2026-05-06-does-not-approximate-attention/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-06-does-not-approximate-attention/</guid>
        <pubDate>Wed, 06 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted main post via api.fxtwitter.com fallback and checked the linked SubQ technical post.
Mario Zechner is skeptical of SubQ’s claim that SSA does not approximate attention; his objection is that unless ignored query-key pairs are provably zero-contribution, selective sparsification is still an approximation.
He also flags the missing detail that really matters: how the model chooses which query-key pairs to keep.
The linked SubQ write-up claims content-dependent selection routes attention only to positions that carry signal, yielding linear scaling …</description>
      </item>
    
      <item>
        <title>dreaming</title>
        <link>https://reading-list.oddship.net/notes/2026-05-06-dreaming/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-06-dreaming/</guid>
        <pubDate>Wed, 06 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Claude Managed Agents update centered on three things: dreaming, outcomes, and multiagent orchestration.
Dreaming is a research-preview async job that reads an existing memory store plus past session transcripts and emits a cleaned&amp;amp;#x2F;reorganized memory store with deduped facts, replaced stale entries, and new synthesized insights; original store remains unchanged.
Outcomes adds an explicit done target plus rubric-driven grading, turning a session from chat into iterative artifact production with a separate grader context feeding gap reports back to the agen…</description>
      </item>
    
      <item>
        <title>HTML5+CSS face lift for the generated pages</title>
        <link>https://reading-list.oddship.net/notes/2026-05-06-html5-css-face-lift-for-the-generated-pages/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-06-html5-css-face-lift-for-the-generated-pages/</guid>
        <pubDate>Wed, 06 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

GitHub PR title: HTML5+CSS face lift for the generated pages by knadh on mitmproxy&amp;amp;#x2F;pdoc; merged Nov 20, 2014.
Logged as a folklore&amp;amp;#x2F;historical reference rather than a current article; likely relevant as an old design&amp;amp;#x2F;implementation artifact in the pdoc&amp;amp;#x2F;docsite lineage.
Retrieval from the public PR page was partial because logged-out GitHub readability extraction is thin, but title&amp;amp;#x2F;author&amp;amp;#x2F;repo&amp;amp;#x2F;merged status were captured.
Follow-up if needed: inspect commits&amp;amp;#x2F;diff directly or use GitHub API&amp;amp;#x2F;source checkout for the substantive changes.

</description>
      </item>
    
      <item>
        <title>Satya&#x2F;Microsoft framing: firms need to redesign work around agentic systems, with AI taking more execution...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-06-satya-microsoft-framing-firms-need-to-redesign-work-around-agentic-systems-with-ai-taking-more-exec/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-06-satya-microsoft-framing-firms-need-to-redesign-work-around-agentic-systems-with-ai-taking-more-exec/</guid>
        <pubDate>Wed, 06 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted main post via api.fxtwitter.com fallback and read the linked Microsoft Work Trend Index piece Agents, human agency, and the opportunity for organizations.
Satya&amp;amp;#x2F;Microsoft framing: firms need to redesign work around agentic systems, with AI taking more execution while humans shift toward judgment, intent-setting, and owning outcomes.
The more interesting claim is organizational, not individual: Microsoft says culture, manager support, and talent practices explain more than 2x the reported AI impact of individual effort alone.
Key vocabulary from …</description>
      </item>
    
      <item>
        <title>transfer station</title>
        <link>https://reading-list.oddship.net/notes/2026-05-06-transfer-station/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-06-transfer-station/</guid>
        <pubDate>Wed, 06 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted main post via api.fxtwitter.com fallback and read linked ChinaTalk piece How to Buy Cheap Claude Tokens in China.
Kyle Chan highlights Zilan Qian’s write-up on the transfer station economy around blocked frontier-model access in China.
Core claim: this is not just a handful of labs evading restrictions, but a broader gray-market stack of intermediaries, payments, proxying, account supply, and abuse adaptation serving ordinary developers, hobbyists, and companies.
Most important insight is governance-related, not the mechanics: each added provide…</description>
      </item>
    
      <item>
        <title>Fragments: May 5</title>
        <link>https://reading-list.oddship.net/notes/2026-05-05-fragments-may-5/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-05-fragments-may-5/</guid>
        <pubDate>Tue, 05 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted via api.fxtwitter.com fallback.
Martin Fowler Fragments: May 5 roundup linking: open-source framework for prompting patterns, musician suing Google for defamation, Apple rethinking AI spend, running LLMs locally, and whether The Genie gets caught in the tar pit.
Link target: https:&amp;amp;#x2F;&amp;amp;#x2F;martinfowler.com&amp;amp;#x2F;fragments&amp;amp;#x2F;2026-05-05.html
Likely useful as a curated bundle rather than a single thesis; good source to revisit for one or two standout downstream links.

Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X track yo…</description>
      </item>
    
      <item>
        <title>Frank&#x2F;jedisct1: SKILL.md is fine for static instructions. But many useful agent workflows are not just inst...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-05-frank-jedisct1-skill-md-is-fine-for-static-instructions-but-many-useful-agent-workflows-are-not-jus/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-05-frank-jedisct1-skill-md-is-fine-for-static-instructions-but-many-useful-agent-workflows-are-not-jus/</guid>
        <pubDate>Tue, 05 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted via api.fxtwitter.com fallback.
Frank&amp;amp;#x2F;jedisct1: SKILL.md is fine for static instructions. But many useful agent workflows are not just instructions. They are loops. Introducing Agent MetaSKILLs.
Linked page: https:&amp;amp;#x2F;&amp;amp;#x2F;swival.dev&amp;amp;#x2F;pages&amp;amp;#x2F;metaskills.html
Relevance to https:&amp;amp;#x2F;&amp;amp;#x2F;rohanverma.net&amp;amp;#x2F;pages&amp;amp;#x2F;harness-engineering&amp;amp;#x2F;: strong fit with the site’s emphasis on harnesses as loops, feedback systems, progressive knowledge, and infrastructure around the model rather than the model alone.
Especially adjacent to sections on Skills, Meta-Skills, The Loop, The Dae…</description>
      </item>
    
      <item>
        <title>Pratilekha</title>
        <link>https://reading-list.oddship.net/notes/2026-05-05-pratilekha/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-05-pratilekha/</guid>
        <pubDate>Tue, 05 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted via api.fxtwitter.com fallback.
Uttaran Nayak (Bangalore) announcing Pratilekha: one API, every Indian &amp;amp;amp;amp; regional language. and we built this ourselves.
Early signal worth tracking as part of the India&amp;amp;#x2F;Bangalore AI&amp;amp;#x2F;app layer scene, especially around multilingual infrastructure rather than generic model wrappers.
Good follow-up question later: what is actually novel here, translation, speech, multilingual inference stack, or developer platform packaging?

Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can let X t…</description>
      </item>
    
      <item>
        <title>Simone&#x2F;evilsocket amplifying claim that Chrome silently installs a 4 GB Gemini Nano model on user devices,...</title>
        <link>https://reading-list.oddship.net/notes/2026-05-05-simone-evilsocket-amplifying-claim-that-chrome-silently-installs-a-4-gb-gemini-nano-model-on-user-d/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-05-simone-evilsocket-amplifying-claim-that-chrome-silently-installs-a-4-gb-gemini-nano-model-on-user-d/</guid>
        <pubDate>Tue, 05 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Extracted via api.fxtwitter.com fallback.
Simone&amp;amp;#x2F;evilsocket amplifying claim that Chrome silently installs a 4 GB Gemini Nano model on user devices, without clear consent prompt, and re-downloads it if deleted.
Linked article: https:&amp;amp;#x2F;&amp;amp;#x2F;awesomeagents.ai&amp;amp;#x2F;news&amp;amp;#x2F;chrome-gemini-nano-silent-install&amp;amp;#x2F;
Why it matters: local&amp;amp;#x2F;on-device AI is increasingly shipping as platform behavior, not just user choice; good angle around consent, storage&amp;amp;#x2F;bandwidth costs, and silent AI infra deployment.

Embedded source

  
    X &amp;amp;#x2F; Twitter post
    Show embedded post
    X embeds can…</description>
      </item>
    
      <item>
        <title>SubQ</title>
        <link>https://reading-list.oddship.net/notes/2026-05-05-subq/</link>
        <guid>https://reading-list.oddship.net/notes/2026-05-05-subq/</guid>
        <pubDate>Tue, 05 May 2026 00:00:00 +0000</pubDate>
        <description>Imported from historical reading log.

Main post successfully extracted via api.fxtwitter.com fallback.
Post by Alexander Whedon introducing SubQ as a sparse-attention LLM architecture claim: fully sub-quadratic sparse attention, 12M token context window, 52x faster than FlashAttention at 1M tokens, under 5% of Opus cost, and nearly 1,000x less compute by focusing only on relationships that matter.
Core framing: standard transformer attention computes many unnecessary token relationships; sparse attention focuses only on the small fraction that matters.
No tweet.article block present in the AP…</description>
      </item>
    
  </channel>
</rss>
