{"id":"fresh-20260623-001","published_utc":"2026-06-22T21:24:12Z","question":"According to NASA's June 22, 2026 SEWP VI contract release, what are the three acquisition categories, what is the ordering period, and what is the maximum value per contract?","answer":"Categories: Category A IT Solutions; Category B Enterprise-wide IT Service Solutions; Category C IT Mission-Based Services. Ordering period: 10 years, Nov. 1 through Oct. 31, 2036. Maximum value: $20 billion per IDIQ contract.","source":"https://www.nasa.gov/news-release/nasa-awards-solutions-for-federal-enterprise-procurement-contracts/","why_vanilla_should_fail":"Official release published within the last ~24 hours; asks for exact procurement details."} {"id":"fresh-20260623-002","published_utc":"2026-06-22T20:42:34Z","question":"For NASA's June 2026 RockSatX/RockOn combined sounding rocket mission, when is the launch window, how many participants/teams are involved, and approximately how many experiments will the rocket carry?","answer":"Launch window: Wednesday, June 24, 2026, 5:30–9:30 a.m. EDT, with backup Thursday, June 25. Nearly 250 participants from 38 university/community college teams; nearly 50 experiments.","source":"https://www.nasa.gov/centers-and-facilities/wallops/nasa-sounding-rocket-to-launch-student-experiments/","why_vanilla_should_fail":"Fresh NASA operational announcement with exact timing and counts."} {"id":"fresh-20260623-003","published_utc":"2026-06-22T19:21:27Z","question":"In NASA's June 22, 2026 media advisory, which country is scheduled to sign the Artemis Accords, at what time/date, who will host, and what signer number will it become?","answer":"Botswana; 9:30 a.m. EDT Thursday, June 25, 2026; hosted by NASA Deputy Administrator Matt Anderson; Botswana will be the 68th country to sign.","source":"https://www.nasa.gov/news-release/nasa-invites-media-to-botswana-artemis-accords-signing-ceremony/","why_vanilla_should_fail":"Fresh diplomatic/media advisory; exact event facts are not parametric knowledge."} {"id":"fresh-20260623-004","published_utc":"2026-06-22T15:00:00Z","question":"According to NASA's June 22, 2026 Webb story on comet 3I/ATLAS, which Webb instrument was used, what unusual chemical measurements were highlighted, and where/when was the paper published?","answer":"Instrument: NIRSpec / Near-Infrared Spectrograph. Measurements: carbon and deuterium/heavy-hydrogen chemical ratios unlike solar-system comets. Paper published June 22 in Nature.","source":"https://science.nasa.gov/missions/webb/nasas-webb-finds-clues-to-ancient-distant-origin-of-comet-3i-atlas/","why_vanilla_should_fail":"Fresh science result; asks for exact instrument and publication facts."} {"id":"fresh-20260623-005","published_utc":"2026-06-22T17:39:37Z","question":"In NASA's June 22, 2026 Chandra image article, where is the possible supernova remnant located and what would make it notable if confirmed?","answer":"It is in the middle/central region of the Milky Way. If confirmed, it would be one of the closest supernova remnants ever discovered to the supermassive black hole at the Galactic Center.","source":"https://www.nasa.gov/image-article/nasas-chandra-finds-possible-supernova-remnant/","why_vanilla_should_fail":"Fresh image article; answer depends on source-specific wording."} {"id":"fresh-20260623-006","published_utc":"2026-06-22T16:45:24Z","question":"For NASA's US Spacewalk 95 announcement, what task will astronauts perform, when is the spacewalk scheduled to begin, and who are the three preview briefing participants listed?","answer":"Task: replace a wrist joint on the ISS Canadarm2 robotic arm. Start: approximately 8:35 a.m. EDT Tuesday, June 30, 2026. Briefing participants: Bill Spetch, Fiona Antkowiak, and Jason Dyer.","source":"https://www.nasa.gov/news-release/nasa-to-cover-us-spacewalk-95-host-preview-news-conference/","why_vanilla_should_fail":"Fresh schedule/personnel details requiring retrieval."} {"id":"fresh-20260623-007","published_utc":"2026-06-22T17:59:55Z","question":"What real-world data-collection bottleneck does the June 22, 2026 arXiv paper 'AutoDex' claim to address, and what loop must run without human intervention?","answer":"It addresses scalable real-world dexterous grasping data collection: teleoperation is slow/operator-biased and simulation cannot certify contact validity. The loop is perception, execution, labeling, and reset running without human intervention.","source":"https://arxiv.org/abs/2606.23689v1","why_vanilla_should_fail":"New arXiv preprint posted June 22; requires abstract-specific claims."} {"id":"fresh-20260623-008","published_utc":"2026-06-22T17:59:53Z","question":"In 'Randomized YaRN Improves Length Generalization for Long-Context Reasoning,' what three components are combined in the proposed training method?","answer":"YaRN-based positional extrapolation, randomized positional encoding, and a length curriculum.","source":"https://arxiv.org/abs/2606.23687v1","why_vanilla_should_fail":"New arXiv preprint with method-specific details."} {"id":"fresh-20260623-009","published_utc":"2026-06-22T17:59:20Z","question":"What stop-and-go simplification does 'CoorDex' criticize, and what control formulation does it introduce?","answer":"It criticizes walking to an object, stopping to manipulate it, then resuming locomotion, often with low-DoF open-close end effectors. It introduces coordinated latent residual control for high-DoF dexterous loco-manipulation on the move.","source":"https://arxiv.org/abs/2606.23680v1","why_vanilla_should_fail":"New robotics preprint; exact contribution not in model memory."} {"id":"fresh-20260623-010","published_utc":"2026-06-22T17:59:17Z","question":"What problem with modern text-to-image models motivates 'Semantic Browsing,' and what user capability does the method aim to provide?","answer":"Strict prompt adherence can collapse samples into a single visual interpretation, reducing meaningful diversity. Semantic Browsing aims to let users navigate controlled, structured diversity through meaningful design choices.","source":"https://arxiv.org/abs/2606.23679v1","why_vanilla_should_fail":"New paper; requires source-specific framing."} {"id":"fresh-20260623-011","published_utc":"2026-06-22T17:58:54Z","question":"According to the AIR arXiv abstract, what limitation of prior interleaved-reasoning/tool-use work does AIR target?","answer":"Prior work focuses mainly on predefined heuristic visual manipulations for vision-perception tasks and is inherently unable to address numerical computation problems; AIR targets adaptive interleaved reasoning with code in MLLMs.","source":"https://arxiv.org/abs/2606.23678v1","why_vanilla_should_fail":"Fresh preprint; answer is abstract-specific."} {"id":"fresh-20260623-012","published_utc":"2026-06-22T17:58:52Z","question":"What open theoretical gap does 'Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?' identify, and which optimizers does it contrast with AdamW?","answer":"It identifies the lack of rigorous convergence theory for AdamW under heavy-tailed stochastic gradient noise in LLM pretraining. It contrasts AdamW with sign-based optimizers such as Lion and Muon, and with AdaGrad.","source":"https://arxiv.org/abs/2606.23676v1","why_vanilla_should_fail":"Fresh open-problem paper; exact contrasts require retrieval."} {"id":"fresh-20260623-013","published_utc":"2026-06-22T17:57:15Z","question":"What limitation in existing mental-health assessment approaches does 'PsyBridge' claim to address?","answer":"Existing approaches rely on isolated screening instruments or data-driven models, lack interpretability and multi-dimensional integration, and focus on individual indicators like depression or anxiety rather than comprehensive explainable decision support.","source":"https://arxiv.org/abs/2606.23673v1","why_vanilla_should_fail":"Fresh preprint; exact limitation statement is source-specific."} {"id":"fresh-20260623-014","published_utc":"2026-06-22T17:57:08Z","question":"In the June 22, 2026 arXiv paper on bit manipulation puzzles, what is the task objective and what LLM failure mode do the authors say traditional methods induce?","answer":"Objective: discover a hidden logical rule transforming input binary strings to outputs, then apply it to unseen inputs. Traditional methods force LLMs to simulate complex boolean logic/arithmetic, leading to hallucinations.","source":"https://arxiv.org/abs/2606.23672v1","why_vanilla_should_fail":"Fresh challenge paper; answer depends on new abstract."} {"id":"fresh-20260623-015","published_utc":"2026-06-22T17:56:30Z","question":"What did 'Can LLMs Reliably Self-Report Adversarial Prefills, and How?' find about models recognizing compromised outputs, and what average intent-claim rate is reported?","answer":"Across ten open-weight instruction-tuned LLMs and four safety benchmarks, no model reliably recognized its own compromised outputs; models claimed intent on prefilled responses at an average rate of 27.3%.","source":"https://arxiv.org/abs/2606.23671v1","why_vanilla_should_fail":"Fresh safety preprint with a specific reported number."} {"id":"fresh-20260623-016","published_utc":"2026-06-22T17:56:25Z","question":"What architectural default does 'Tapered Language Models' question, and what asymmetry motivates the question?","answer":"It questions the default stack of identical layers with parameters allocated uniformly across depth. The motivation is evidence that layers contribute non-uniformly, with later layers refining rather than transforming the residual stream.","source":"https://arxiv.org/abs/2606.23670v1","why_vanilla_should_fail":"Fresh architecture preprint; details are not stable prior knowledge."} {"id":"fresh-20260623-017","published_utc":"2026-06-22T17:52:59Z","question":"How does 'On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners' model user-system interaction, and what conceptual decomposition does it introduce?","answer":"It models user-system interaction as a bilevel cheap-talk game. It introduces a decomposition separating task inference from execution.","source":"https://arxiv.org/abs/2606.23668v1","why_vanilla_should_fail":"Fresh theoretical preprint; asks for exact framing."} {"id":"fresh-20260623-018","published_utc":"2026-06-22T17:48:40Z","question":"What does MAS-PromptBench study, and why are system prompts described as an accessible optimization surface in multi-agent systems?","answer":"It studies when prompt optimization improves multi-agent LLM systems. System prompts are accessible because they specify agents' roles/behaviors and can improve the system without model fine-tuning.","source":"https://arxiv.org/abs/2606.23664v1","why_vanilla_should_fail":"Fresh multi-agent benchmark preprint; source-specific."} {"id":"fresh-20260623-019","published_utc":"2026-06-22T00:00:00Z","question":"In Google's June 22, 2026 Jules post, what gap in SWE-Bench-style evaluation is identified, and what is 'insight policy'?","answer":"SWE-Bench evaluates task completion for narrowly defined bugs, but not open-ended goals for proactive agents. Insight policy is the ability to decide what matters, what evidence supports it, and whether to interrupt the developer or stay silent.","source":"https://developers.googleblog.com/measuring-what-matters-with-jules/","why_vanilla_should_fail":"Fresh blog post on agentic coding evaluation; exact term definition requires retrieval."} {"id":"fresh-20260623-020","published_utc":"2026-06-23T03:45:33Z","question":"From the LangChain GitHub release feed around June 22-23, 2026, which four package release tags appeared most recently?","answer":"langchain-openrouter==0.2.4, langchain-openai==1.3.3, langchain-anthropic==1.4.7, and langchain==1.3.11.","source":"https://github.com/langchain-ai/langchain/releases.atom","why_vanilla_should_fail":"Release-feed fact published within hours; exact package versions are unavailable to parametric memory."}