Validating LLM Safety Analysers Beyond STPA Outputs
Why LLM safety analysers themselves must be validated, and what constitutional meta-STPA changes for assurance.
Why LLM safety analysers themselves must be validated, and what constitutional meta-STPA changes for assurance.
Why LLM agreement can mislead evaluation, with correlated errors, shared wrong answers, and safer judging protocols.
A curated link roundup from recently collected official updates and tech news.
Meta’s planned AI chip production from September highlights tighter control over training and inference infrastructure, not just models.
Key issues in the MiniMax report: a rumored 2.7 trillion-parameter LLM, possible open weights, licensing, and inference costs.
RAID found six scoring exploits in NHL 26 goalie AI in one run, highlighting automated QA and reusable red-team testing.
SPEAR links Unreal Engine with Python, targeting 73 fps rendering and 14K+ exposed functions for research workflows.
How contextual inputs and shared recurrence aim to control diverse robot morphologies with one policy across zero-shot and sim-to-real tests.
A look at interpreting transformer-based VLM adversarial vulnerability through intermediate spectral subspaces.
Examines whether closed government-company talks are enough to judge frontier AI release safety and accountability gaps.
A curated link roundup from recently collected official updates and tech news.
A look at an arXiv paper proposing continual learning for adaptive control of modular soft robots under morphology changes.
Gimitest is an open-source framework for testing RL policies under changing conditions to uncover failures and vulnerabilities.
Why agentic AI governance must cover autonomy, tool use, external actions, audit logs, and human oversight.
HIVE evaluates how vision-language hallucinations propagate into later reasoning and distort downstream predictions.
An overview of PCBWorld, a KiCad-based environment for evaluating PCB routing AI with native actions and DRC feedback.
Examines whether reusable skill files improve quality, auditability, and operations in repetitive AI data science tasks.
VASP Agent targets reliable scientific automation by combining input consistency, long-run supervision, and output validation.
Examines how conversational AI and games compete for attention, highlighting different user needs and social dynamics.
Why agent safety must verify execution, tool use, and state changes, not just final responses.
Examines how attention-limited pairwise labels in RLHF can distort reward learning and be mistaken for true preference.
A look at training small models to find first reasoning errors, use structured feedback, and revise answers in physics tasks.
Why long-video AI struggles with narrative and causal links, and how hierarchical memory and agentic reasoning help.
Explains why better LLM performance and office automation do not directly reduce electricity, rent, or food costs.