THREAT OPS › Threat News › Sol Searching | Can Frontier Models Tackle Autonomous Long-Horizon Malware Analysis?
Sol Searching | Can Frontier Models Tackle Autonomous Long-Horizon Malware Analysis?
<h2>Executive Summary</h2> <ul> <li>SentinelLABS developed a multi-stage reverse-engineering benchmark for the latest generation of frontier models by recreating our recent investigation of <a href="https://s1.ai/fast16" rel="noopener noreferrer" target="_blank">fast16</a>, a unique 2005 sabotage implant.</li> <li>Most AI benchmarks test bounded tasks. This benchmark tests whether a model can keep
MITRE ATT&CK techniques
- ServerlessT1583.007
- Artificial IntelligenceT1588.007
- Windows ServiceT1543.003
- ServerlessT1584.007
- ServerlessAML.T0008.004
Indicators of compromise
- CVE-2025-37899cve
- https://s1.ai/fast16url
- https://alperovitch.sais.jhu.edu/an-experiment-in-malware-reverse-engineering/url
- https://sean.heelan.io/2025/05/22/how-i-used-o3-to-find-cve-2025-37899-a-remote-zeroday-vulnerability-in-the-linux-kernels-smb-implementation/url
- https://openai.com/daybreak/url
- https://www.anthropic.com/project/glasswingurl
- https://blackhat.com/asia-26/briefings/schedule/#cyber-paleontology-in-the-age-of-ai-51494url
- https://nono.sh/url
- https://open.spotify.com/episode/1Ii8doMtZt7CTLU6ABgif5?si=MxKaZx0eTJe2RX7KTUR_wAurl
- https://fireworks.ai/url