Community journal

The latest from the MySQL community

Ideas, releases, practical guides, and perspectives from the people building with MySQL.

Follow the feed
Showing entries 1 to 3 of 3 « Previous | Next »
Displaying posts with tag: Compiler Optimization (reset)
PGO or not PGO this is the dilemma. Step 3

Step 3: Why my sysbench-trained build loses and how to do it right.

Given the topic complexity and the length of this article I have split it in 3 three different blog-post:

  1. What is PGO
  2. How PGO it works
  3. Why my sysbench-trained build loses and how to do it right.

Three compounding reasons:

  1. Uncovered code gets pessimized.
    Sysbench-tpcc touches a narrow slice of mysqld.
    Every function with zero counts is treated as cold: GCC optimizes it for size, skips inlining, and shoves it into cold sections.
    But at runtime I still execute plenty of code my training never touched, such as purge, flushing, stats …
[Read more]
PGO or not PGO this is the dilemma. Step 2

Step 2: How it works

Given the topic complexity and the length of this article I have split it in 3 three different blog-post:

  1. What is PGO
  2. How PGO it works
  3. Why my sysbench-trained build loses and how to do it right.

How PGO works: PGO is a two-pass build. First pass compiles with instrumentation (-fprofile-generate): every basic block and branch gets a counter.
You run a training workload, counters are dumped to profraw files.
Second pass recompiles using those counts to drive inlining decisions, branch layout (hot path falls through, cold path jumps away), hot/cold function splitting, code ordering for icache/iTLB locality, …

[Read more]
PGO or not PGO this is the dilemma. Step 1

Given the topic complexity and the length of this article I have split it in 3 three different blog-post:

  1. What is PGO
  2. How PGO it works
  3. Why my sysbench-trained build loses and how to do it right.

Step 1: What is PGO

PGO (Profile-Guided Optimization) is a two-pass compilation technique. 

First, we build the program with instrumentation that adds counters to every branch, loop, and call site.
Next, we run it against a representative training workload so these counters can record which code paths are actually hot versus cold.
We then recompile the program using that data to drive the compiler’s decisions.

With this profiling …

[Read more]