MLSecRecommendations
Сигналы по MLSec

Лента новостей —
что происходит в безопасности ML и LLM прямо сейчас

Стандарты, инструменты, инциденты, релизы. 13 активных источников: OWASP, MITRE ATLAS, NIST, NCSC, лаборатории Anthropic / Google / Microsoft, профильные вендоры и тематические Telegram-каналы. Бесплатно и без регистрации.

8350 в выдаче13 источниковобновлено 2 сент., 19:46 RSS-фид
8 350 результатов · 30 непрочитано
Вчера30
en
My coding agents kept saying they were done when they had only finished typing

I kept running into a boring failure with coding agents: the completion report sounded much more reliable than the work behind it. Sometimes the agent ran a test, but not the one that covered the change. Sometimes it misread the output. Sometimes it claimed the check passed witho

en
What are the must-checks when shopping for LLM Gateways?

Been spending some time learning about this piece of fascinating technology lately and I found that it can help me with some of my orgs problems. Currently serving 400+ people through bedrock and I am interested in adding a gateway layer to have some: - rbac capabilities for MCPs

ru
«Спорт, борщ, крипта»: как мы проверяли, работают ли интересы в рекомендательной системе

Пользователь пишет интересы текстом, мы делаем из них один эмбеддинг. Работает он как монета (ROC-AUC 0.52), а на коротких интересах 0.41, то есть хуже случайного порядка. Прогнали шесть вариантов починки и подняли до 0.80. Рассказываю, что сработало, а что нет. Что показали цифр

ru
«Спорт, борщ, крипта»: как мы проверяли, работают ли интересы в рекомендательной системе

Пользователь пишет интересы текстом, мы делаем из них один эмбеддинг. Работает он как монета (ROC-AUC 0.52), а на коротких интересах 0.41, то есть хуже случайного порядка. Прогнали шесть вариантов починки и подняли до 0.80. Рассказываю, что сработало, а что нет. Что показали цифр

en
OpenLeash Adds a Human Check to Risky AI Agent Actions

The security tool intercepts potentially dangerous agent actions, blocking clear threats and requesting human approval when intent is uncertain. The post OpenLeash Adds a Human Check to Risky AI Agent Actions appeared first on SecurityWeek.

en
Model injection: making the model the payload for attacks

This is my second post here! Cool if you check it out! Short version: i built a POC where a fine-tuned open model behaves as it should until it sees a specific trigger, then drops its guardrails and instructions. the keyword version is a toy, but the real point is that the trigge

en
AI Agents Are Now Emailing Me with Their Security Concerns

I received the two emails below earlier in the month. They’re vaguely coherent. I suppose I shouldn’t be surprised that the corpus that AIs are training on contain data suggesting that I am someone to write to with random computer and network security problems. After all, I obser

en
Have u used Genie code to build LLM apps ?

I was curious if anyone is using Genie Code beyond usual SQL generation. for ex. to build or iterate on agent workflow or tool integrations. Was wondering how useful it is in a real LLM app development workflow. Iam asking because started to use it for an LLM app and want to see

ru
Как подключить к LLM вашу документацию и базу знаний

LLM умеет отвечать на вопросы, но не знает особенностей вашего проекта и внутренней документации. В статье разберём, как с помощью RAG подключить собственные данные к модели: от подготовки документов и эмбеддингов до поиска нужного контекста и генерации ответов. Разобрать подход

en
UK Moves to Block High-Risk Tech Suppliers From Critical Infrastructure

Late amendments to the Cyber Security and Resilience Bill would give ministers new powers to restrict risky technology providers as supply chain attacks intensify. The post UK Moves to Block High-Risk Tech Suppliers From Critical Infrastructure appeared first on SecurityWeek.

ru
Бенчмарки Мебиуса: зачем IBM учит ИИ проверять собственные вопросы

Для оценки эффективности ИИ уже давно и успешно используют самые разные бенчмарки. Берем модель/агента, выдаем им одинаковый набор задач, считаем долю успешных решений и получаем удобную цифру, по которой можно сравнивать системы между собой.По результатам используем ту, которая

ru
Бенчмарки Мебиуса: зачем IBM учит ИИ проверять собственные вопросы

Для оценки эффективности ИИ уже давно и успешно используют самые разные бенчмарки. Берем модель/агента, выдаем им одинаковый набор задач, считаем долю успешных решений и получаем удобную цифру, по которой можно сравнивать системы между собой.По результатам используем ту, которая

en
Claude's new system prompt really doesn't want to reproduce song lyrics

Anthropic publish the system prompts for their Claude consumer applications (Claude.ai and the Claude mobile apps - sadly not for Claude Cowork or Claude Code). I love that they do this, and that they share not just the current prompts but historic changes to their prompts as wel

en
Keep getting infinite loop on Luna 5.6

Hi, we're using Luna api for our agent, in some cases it starts spitting repeating stuff, like tool calls, or hashes... like 10k of them... until we cut it off... does anyone else has the same issue? submitted by /u/spider853 [link] [comments]

en
Meta Ads Push StreamRat Android Trojan That Can Gain Near-Complete Device Control

Cybersecurity researchers have disclosed details of a new Android banking trojan called StreamRat that was promoted to Spanish-speaking users through a fake television-streaming campaign on Meta and can give operators near-complete control of infected devices. ThreatFabric said t

en
Challenges with LLM Quota Increase on Cloud Providers

Anybody here has experience with cloud providers to increase your OpenAI and Anthropic model endpoint quota (TPM) as a startup? I think we are pretty lucky that our product has got more traction than we anticipated, and now it's causing us problems. When usage spikes, we easily h

en
Most open-source AI detectors can't hold a 0.5% false-positive rate [P]

We needed to know where the open-source AI-detection field actually stands, so we ran every notable open detector through the same protocol. Setup: - Public data only: Jabarian & Imas 2025 (NBER), Liang 2023 TOEFL essays, a 1,060-text frontier set (GPT-5.x, Claude Opus 5, Gemini

en
Communicating Under Pressure: Best Practices for Service Providers

Developed by CISA, the Federal Bureau of Investigation, and international partners, this guidance describes how organizations can plan and execute clear, timely, accurate, and audience-appropriate communications during IT and operational technology (OT) outages. Whether caused by

en
CISA Adds Seven Known Exploited Vulnerabilities to Catalog

CISA has added seven new vulnerabilities to its Known Exploited Vulnerabilities (KEV) Catalog, based on evidence of active exploitation. CVE-2026-9586 Sangoma Switchvox SQL Injection Vulnerability CVE-2026-48710 Kludex Starlette HTTP Request/Response Smuggling Vulnerability CVE-2

ru
Qwen заняла первое место с пересекающимися интервалами

2 сентября Qwen3.8-Max-0902 поднялась на первую строку общего зачёта Code Arena WebDev. На момент исходной новости модель набрала 1691 очко и опережала Claude Opus 5 Max на три пункта.Проблема скрывается рядом с красивой цифрой: результат Qwen помечен как предварительный, голосов

Подключите ленту в RSS-клиент: RSS