Skip to content
← All posts
· Delta1 Labs Decompilation.NETDeep Dive

How .NET metadata tables work — and how a decompiler reads them

Every .NET assembly carries a small relational database of its own structure: the metadata tables. Here is what TypeDef, MethodDef and the coded indexes actually contain, how a metadata token addresses a row, and how a decompiler walks the tables to turn a DLL back into types and methods.

Open a .NET DLL in a hex editor and most of it is not code. Before the IL, before the resources, sits a compact relational database describing the assembly’s own structure: its types, methods, fields, the external members it calls, the strings and signatures it uses. This is the metadata, and it is the first thing the runtime reads, the first thing a decompiler reads, and the reason a .NET assembly can be taken apart so cleanly. Understanding the tables explains what a tool like Glass.NET is actually doing when it turns a binary back into types and methods.

A database inside the assembly

The metadata lives in a set of tables, each with a fixed schema defined by ECMA-335. Every table is a list of rows, and every row is a tuple of columns. A column holds either a small integer, an index into a heap (where variable-length data like names and signatures live), or an index into another table. That last kind is what makes metadata relational: rows point at rows.

The heaps are four streams that sit beside the tables:

  • #Strings — all identifier text (type names, method names, namespaces), each a UTF-8 run.
  • #Blob — binary data, most importantly signatures (a method’s parameter and return types are a blob).
  • #GUID — module identifiers.
  • #US — user strings, the literals your code loads with ldstr.

A column never stores a name directly; it stores an index into #Strings. This is why renaming during obfuscation is cheap — only the heap text changes, not the table structure.

The tables most relevant to reconstruction:

TableIdOne row per
Module0x00the module itself
TypeRef0x01a referenced (external) type
TypeDef0x02a type defined here
Field0x04a field
MethodDef0x06a method defined here
MemberRef0x0Aa referenced external member
TypeSpec0x1Ban instantiated generic type

What a TypeDef row holds

A TypeDef row is the anchor for a type. Its columns are:

TypeDef:
  Flags      : UInt32          // visibility, abstract/sealed, layout
  Name       : #Strings index  // "InvoiceService"
  Namespace  : #Strings index  // "Billing"
  Extends    : TypeDefOrRef     // base type (a coded index)
  FieldList  : Field  RID       // first field; runs to next row's FieldList
  MethodList : MethodDef RID    // first method; same run trick

Two ideas do a lot of work here. First, Name and Namespace are heap indexes, so the human-readable identity is one dereference away. Second, FieldList and MethodList are run-encoded: a type’s methods are the MethodDef rows from this row’s MethodList up to (but not including) the next TypeDef’s MethodList. There is no explicit count — the boundary is the next type’s start. A reader that ignores this and treats MethodList as a single method gets everything wrong after the first type, which is a classic first-time-parser bug.

Coded indexes: pointing at “one of several tables”

Extends can be a type defined in this assembly (TypeDef), a type referenced from another (TypeRef), or a TypeSpec. The column must therefore point into one of several tables. Metadata solves this with a coded index: a few low bits select which table, and the rest is the RID. For TypeDefOrRef, 2 bits choose among TypeDef/TypeRef/TypeSpec and the remaining bits are the row. So decoding Extends means masking off the tag bits to learn the table, then shifting to get the row. Coded indexes are everywhere in metadata and are the main reason a naive “read it as a flat struct” approach fails; you must decode each column per its defined type.

Tokens: how IL names a row

When IL needs to refer to a method, type or field, it uses a metadata token — a 4-byte value where the high byte is the table id and the low three bytes are the RID (1-based):

0x060x00 00 04table idrow id (RID)MethodDefrow 4token = 0x06000004

So 0x02000003 is “TypeDef, row 3”; 0x0A000011 is “MemberRef, row 17.” A call instruction’s operand is exactly such a token. To know what call 0x0A000011 invokes, the reader jumps to MemberRef row 17, reads its Class (what the member is on) and Name/Signature columns, and resolves them through the heaps and other tables. This single mechanism — token in, row out — is how all the cross-references in IL are followed.

Walking the tables to reconstruct a method

Put it together and the decompiler’s structural pass is a sequence of table walks:

// Pseudocode over the raw tables (names/sigs via the heaps):
foreach (var typeRow in tables.TypeDef)              // 0x02
{
    var typeName = strings[typeRow.Name];            // #Strings
    var ns       = strings[typeRow.Namespace];
    var baseType = Resolve(typeRow.Extends);         // decode coded index

    foreach (var m in MethodsOf(typeRow))            // run-encoded MethodList
    {
        var name = strings[m.Name];
        var sig  = blob[m.Signature];                // #Blob → params/return
        var il   = ReadBodyAt(m.Rva);                // RVA → IL bytes
        // tokens inside `il` resolve back through TypeRef/MemberRef/TypeDef
    }
}

ReadBodyAt(m.Rva) is the bridge from metadata to code: the MethodDef row’s RVA column is a relative virtual address pointing at the method’s IL body in the PE file. The body has a tiny header (giving the code size and, for fat headers, a token for the local-variable signature), then the raw IL. The decompiler decodes that IL, and every operand token sends it back into the tables. Names come from #Strings, type shapes from the #Blob signatures. Only after this structural recovery does the harder, heuristic work begin — turning verified IL and metadata into readable C#, rebuilding loops, using blocks and the rest.

Why this matters when you read a binary

Because the structure is a standardized database, a decompiler can show you a type’s exact shape — members, base type, signatures, attributes — with certainty, even when the body is hard to reconstruct or deliberately protected. It also explains what obfuscation can and cannot touch: renaming rewrites #Strings entries but leaves the table graph intact, so the structure is still fully readable; it is the IL bodies and control flow that stronger protection reshapes. When you open an assembly in Glass.NET and drill from a type to its methods to the raw IL, you are walking exactly these tables — TypeDef to MethodDef to RVA to IL, resolving tokens the whole way down. The metadata is the map; the decompiler just reads it out loud.

Try Nebula.NET

Harden your .NET code in minutes — start with the free edition.