Research Theme

Security

We study how to find weaknesses in binaries without source code, applications running on TEEs, and email-sending infrastructure, and how far those findings can be trusted. Our work spans both LLM-based analysis and Internet-scale measurement.

LLM decompilationTEE taint analysisEmail & DNS measurement

What we work on

  • Recover C code with LLMs from IoT firmware and binaries whose source code is unavailable, and inspect the quality of the result.
  • Use LLMs to trace data flowing across trust boundaries in TEE applications and detect design flaws.
  • Measure reverse DNS of email senders at Internet scale and test whether it is a valid indicator for phishing detection.
  • Rather than stopping at "we detected it," we focus on showing how far each indicator can be trusted.

Papers

Security / Reverse Engineering

Reading binaries with LLMs

Model-Agnostic Staged Pipeline for LLM Decompilation with Quality Inspection

Targeting IoT firmware and executables whose source code no longer exists, this study uses LLMs to recover C code and inspects, step by step, how far the result can be trusted.

LLMDecompilationIoT firmwareVulnerability analysis
Show details Hide details

Problem

Conventional LLM decompilation often takes the generated code as a single output, making it hard to see the gap between code that compiles and code that behaves like the original.

Approach

We built a pipeline that iterates generation and repair, combining Ghidra pseudo-C, function signatures, rule-based quality checks, compiler diagnostics, and execution results.

Key results

Evaluated on HumanEval, MBPP, ExeBench, and OpenWrt, the pipeline improved the recompilation rate on every dataset. We also show that successful recompilation does not guarantee matching behavior.

Security / Static Analysis

Finding boundary mistakes in TEE apps

LLM-Driven Taint Analysis for Detecting Bad Partitioning Issues in Trusted Applications

This study uses LLM-driven taint analysis to detect trust-boundary design mistakes in Trusted Applications running in a Trusted Execution Environment.

TEETrusted ApplicationTaint analysisLLM security
Show details Hide details

Problem

A TEE handles data passing between the secure world and the normal execution environment, so mishandling shared memory or input validation can lead to leaks of sensitive data or unauthorized memory operations.

Approach

We extract a broad set of candidate functions and trace call flows in three stages: START, MIDDLE, and END. Knowledge of TEE-specific APIs and trust boundaries is supplied to the prompt separately.

Key results

The approach detected issues at different locations from the existing rule-based analyzer DITING, and combining the two improved recall. This points toward using it as a complementary analysis while keeping false positives low.

Security / Internet Measurement

Measuring the trustworthiness of email senders

Characterizing Reverse DNS Deployment for SPF-indicated Sender IP Addresses

This study examines how reverse DNS, which is relevant to anti-phishing measures, is configured across a large set of candidate sender IPs extracted from SPF records.

DNSSPFFCrDNSAnti-phishing
Show details Hide details

Problem

The presence of PTR records or FCrDNS is used as a clue for spotting suspicious senders, but legitimate senders sometimes lack them, so relying on it alone risks false positives.

Approach

We collected SPF records from the Tranco Top 1M domains and classified the PTR and FCrDNS status of more than 20.96 million candidate IPv4 sender addresses.

Key results

Only 41.1% had a PTR record and only 23.0% passed FCrDNS. Reverse DNS is a useful signal, but it should be read together with authentication, reputation, and behavior.

Ongoing Research

Research topics the lab is working on as of the 2026 academic year.

Improving the accuracy of LLM-based decompilation

We improve recovery accuracy with a pipeline in which LLMs generate test cases and C code, which are then verified by recompiling and re-running them. We are also looking for students who would like to help with this research.

Background

Decompilation recovers C code from binaries or assembly and is essential for finding vulnerabilities in IoT devices. Conventional tools such as Ghidra sometimes fail partway through recovery.

LLMDecompilationIoT

A survey of PTR records on mail servers for anti-phishing

We investigate whether reverse DNS configuration helps distinguish legitimate emails from phishing emails, and consider how it can be used.

Background

It is known that many IP addresses sending phishing emails have no reverse DNS (PTR) record configured.

PhishingDNSPTR

A survey of BIMI records in industries targeted by phishing

Focusing on industries in Japan that are frequent phishing targets, we investigate how BIMI and DMARC are configured and how effective they are, and analyze the current state of anti-phishing measures.

Background

BIMI displays the sending organization's logo on emails that pass SPF, DKIM, and DMARC authentication, and is drawing attention as a way to visually identify legitimate emails.

BIMIDMARCPhishing

Security research for smart homes

We work on using LLMs to analyze TEE programs and detect implementation mistakes that lead to information leaks, and on verifying design gaps in bridges that connect different devices by observing how they actually behave.

Background

In a smart home, even if the app itself is secure, there is no guarantee that the bridges connecting devices or the programs handling sensitive functions are secure too.

IoTTEESmart homes

Prioritizing vulnerability alerts with usage-evidence SBOMs

We automatically determine whether each component is actually used, through source code analysis and runtime observation, and attach that evidence to the SBOM. Based on usage, vulnerability alerts are sorted into "act now," "can wait," and "needs review."

Background

An SBOM is a "bill of materials" recording the names, versions, and dependencies of the packages that make up a piece of software, and is used to check whether it contains vulnerable components.

SBOMVulnerability managementSupply chain

Externally verifying unauthorized voice training in voice-generation AI

Without looking inside the AI, we quantify generated speech and the person's real voice with a speaker recognition model, and statistically detect traces of training from the difference in similarity compared with unrelated voices. We are researching this as an auditing technique.

Background

Voice actors' voices are being used to train AI without permission, producing voices nearly identical to theirs, yet there is still no technology to verify from the outside whether a voice was used for training.

Voice-generation AIRights protectionAI auditing

Contact

+81-47-469-5709

matsuno.yutaka(at)nihon-u.ac.jp

* To prevent spam, "@" is written as "(at)". Please replace (at) with @ when sending.

Matsuno Lab, Department of Computer Engineering, College of Science and Technology, Nihon University

Room 243, 4F, Building 2, 7-24-1 Narashinodai, Funabashi, Chiba 274-8501, Japan

Directions