Why Alignment Shapes LLM Behavior More Than Personality
Apologies, refusals, and sycophancy in LLMs are shaped more by alignment, rewards, and prompting than personality.
Apologies, refusals, and sycophancy in LLMs are shaped more by alignment, rewards, and prompting than personality.
MKGR combines one sequence modality and four knowledge graphs to improve cold-start PPI prediction over prior baselines.
As multiple-choice medical benchmarks saturate, open-ended clinical reasoning and safety are becoming key measures.
Open-weight LLM safety should be judged not only at release, but by how easily fine-tuning can weaken safeguards later.
PACE examines whether low-cost non-agent benchmarks can predict expensive agent benchmark performance.
ReContext highlights that long-context value depends on reusing evidence already in the prompt, not just larger windows.
Why scientific ML paper reproduction needs workflow, progress tracking, and evidence-claim matching beyond code generation.
A summary of arXiv 2607.01793 on automating agent safety testing from risk discovery to evidence-grounded verification.
A paper on combining RLVR with human demonstrations to train style, structure, and diversity beyond verifiable rewards.
Code model evaluation should weigh real task success, retries, latency, and token cost, not benchmark scores alone.
DiscoLoop explores multi-hop reasoning inside a single forward pass without relying on long external CoT tokens.
Examines whether combining rail crossing images with accident records improves safety assessment and what validation matters.
OCB evaluates native Office file understanding, revealing document AI limits beyond PDF-based QA.
AI data center risks hinge less on hype than on grid connection, cooling design, water tracking, and permitting.
DART-VLN targets stale memory reads and local backtracking in discrete VLN using training-free test-time control.
Examines why remote robots on NTN need memory-based communication that uses past link states and task context.
A look at high-risk AI oversight through the humans-as-handlers approach, focusing on intervention, accountability, and trust.
Examines whether AI safety remains consistent in long conversations and highlights gaps in session-level evaluation.
Examines whether AI eliminates jobs or redesigns tasks, and why this shift matters for hiring, reskilling, and productivity.
A practical guide to balancing agent autonomy, traceability, and control in enterprise orchestration design.
EM may depend on optimizers and batch settings, making finetuning recipes part of safety evaluation, not just data.
Why AI agents must move beyond preference elicitation to support preference formation, with evaluation and safety in view.
How a dual-agent LLM pipeline separates proposing tighter relaxations from verification in automated research.
In education, AI design matters more than raw performance, with student privacy, data minimization, and teacher control at stake.