Home - Waterfall Grid T-Grid Console Builders Recent Builds Buildslaves Changesources - JSON API - About

Console View


Categories: connectors experimental galera main
Legend:   Passed Failed Warnings Failed Again Running Exception Offline No data

connectors experimental galera main
bsrikanth-mariadb
MDEV-40837 Fix crash recording opt context with NULL character_set_results

Unlike character_set_client or collation_connection,
character_set_results can be set to NULL. When recording context for
a query in opt_context_store_replay.cc,
character_set_results->cs_name was accessed without first checking
whether character_set_results itself was NULL, causing a crash.

Fix Optimizer_context_recorder::dump_sql_script() to check
character_set_results for NULL before accessing cs_name, writing
"SET character_set_results=NULL;" to the replay script in that case
so replay reproduces the recorded session instead of silently
defaulting to the script's SET NAMES charset.

Added a test for the same, in opt_context_store_stats.test, that
also checks the recorded script contains the NULL SET statement.
Sergei Golubchik
MDEV-40336 SET AUTHORIZATION un-expired passwords

don't allow SET SESSION AUTHORIZATION if the password expired
Georgi (Joro) Kodinov
MDEV-39307: Remove the verbose part of a comment.
Dave Gosselin
MDEV-41211:  FederatedX multi-table DELETE keeps a const table row

A multi-table DELETE on a FederatedX table left a row in place when a
primary key lookup made the target a const table.  The optimizer reads
a const table's row through index_read_idx_map(), whose default
implementation ends the index scan and frees the result set.  The
server asks for the row's position later, during execution, so the
saved position was empty and the row was skipped.

FederatedX now overrides index_read_idx_map() so that the lookup leaves
its result set open, as index_read() does.  position() then records a
valid position, and the result set is freed at the end of the
statement.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

In progress
Marko Mäkelä
MDEV-41310: Assertion current_lsn < archive_header_was_reset failed

log_t::write_checkpoint(): Correct an off-by-one error in the assertion
expression. We may reach the very end of the current log file.
A subsequent log_t::write_buf() will invoke archive_new_write(),
which will create a new log file if needed.

Tested by: Saahil Alam
Thirunarayanan Balathandayuthapani
MDEV-40274 Server aborts while reloading the table after failed DISCARD TABLESPACE

Problem:
========
row_discard_tablespace() updates SYS_TABLES, SYS_INDEXES and
reassigns the table identifier in SYSTEM TABLES.
These changes cannot be rolled back, and
row_discard_tablespace_for_mysql() commits the transaction even when
row_discard_tablespace() returned an error.
If the reassignment in row_mysql_table_id_reassign() fails
in the middle then the committed data dictionary is inconsistent
and the table can no longer be loaded from it, while the
in-memory table definition stays in the dictionary cache and
remains usable.

Solution:
=========
row_discard_tablespace_for_mysql(): If row_discard_tablespace()
failed, flag the table as corrupted and report the data dictionary
inconsistency in the error log.

dict_table_open_on_name(): Check dict_table_t::space before reading
fil_space_t::get_compression_algo() from it.
Alexey (Holyfoot) Botchkov
MDEV-39923 HANDLER read crashes restore_position() for InnoDB
partitioned table.

m_top_entry has to be reset in ha_partition::index_init().
A stale value left over from an earlier ordered scan on this handler
breaks the execution.
Aleksey Midenkov
MDEV-41344 Timestamp conversion does not work for TRADITIONAL mode

Field_temporal::get_copy_func() returns do_field_datetime whenever the
session's sql_mode has NO_ZERO_DATE or NO_ZERO_IN_DATE (e.g. under
sql_mode=TRADITIONAL), or when the two fields are not eq_def(),
regardless of whether the field is a system-versioned
row_end.

get_copy_func() checked for that value and returned early, before ever
reaching the VERS_ROW_END check that installs
do_field_versioned_timestamp. So an ALTER TABLE .. FORCE meant to
convert an old-format row_end silently copied the value unchanged
instead: the conversion check correctly demanded a copy, but the copy
step never applied it, with no error or warning.

The fix checks the row_end conversion need first, independently of
what Field_temporal::get_copy_func() picked. row_end is
server-maintained and never zero, so NO_ZERO_DATE does not apply to
it.

Tested by "traditional" combination in old_timestamp.test, the test
restores real pre-11.5 row_end fixtures and runs mariadb-upgrade
--force.
Yuchen Pei
MDEV-41322 Convert KEY_NOT_FOUND to EOF in pq-less partition index scans during index_next/index_prev calls

ha_partition::handle_unordered_scan_next_partition is called in a
variety of accesses, including index_read, index_prev, and index_next.

ha_partition::handle_unordered_next and
ha_partition::handle_unordered_prev are called from index_next and
index_prev accesses. They check for signs of end of
scan (m_part_spec.start_part == NO_CURRENT_PART_ID) and exit early if
that is the case. HA_ERR_END_OF_FILE (EOF) is commonly paired with
assigning NO_CURRENT_PART_ID to m_part_spec.start_part whenever
possible, to signal the end of scan.

The error HA_ERR_KEY_NOT_FOUND means the requested key is not found.
It should not mean the end of scan, when for example
ha_partition::handle_unordered_scan_next_partition is called from
index_read, because a subsequent index_next / index_prev call would
then incorrectly return immediately from
ha_partition::handle_unordered_next /
ha_partition::handle_unordered_prev. This is why HA_ERR_KEY_NOT_FOUND
is retained and returned in
ha_partition::handle_unordered_scan_next_partition.

But if the call is from index_next / index_prev, HA_ERR_KEY_NOT_FOUND
should indeed mean end of scan. In this patch, we ensure this is the
case by converting HA_ERR_KEY_NOT_FOUND to EOF in
ha_partition::handle_unordered_next /
ha_partition::handle_unordered_prev.
Kristian Nielsen
MDEV-41264: binlog-in-engine dump thread not stopping at server shutdown

The binlog-in-engine code for binlog dump thread didn't check
should_stop(info, true) to stop on KILL_SERVER when it reaches the end of
the current binlog. This caused the dump threads to hang indefinitely during
SHUTDOWN WAIT FOR ALL SLAVES, and thus the shutdown to hang as well.

Signed-off-by: Kristian Nielsen <[email protected]>
Khaled Riyad
MDEV-40551 Copy/Paste friendly output format for MariaDB Command Line Client

Copy/paste friendly output was only reachable by starting the client with
--silent --skip-column-names, which cannot be done from a running
interactive session.

Add \S, a statement terminator which prints the result of one statement in
the tab separated format without column names.

com_silent() sets output_plain, opt_silent and column_names around
com_go(), then restores them, the same way com_ego() handles vertical.
output_plain selects print_tab_data() ahead of the vertical and table
branches, so \S gives the same output whether the session was started
plainly or with --table, --vertical or --silent. --html and --xml still
win, matching \G.
DerZc
MDEV-40820 Wrong results: SELECT DISTINCT / GROUP BY returns duplicate rows on a RANGE-partitioned table when served by a covering index range scan

SELECT DISTINCT or GROUP BY can return duplicate values from a RANGE-
partitioned table when a covering index range scan incorrectly bypasses
the merge of partition scans.

ha_partition::can_skip_merging_scans() checks only the current multi-
range prefix. Later ranges can have different prefix values, so the
partition outputs do not have the ordering required to skip the
priority-queue merge.

Check every multi-range entry before bypassing the partition-scan merge.
Require both endpoints to bind the complete unordered prefix and to
agree on its bytes. Require that prefix to be the same across all
ranges; otherwise keep the normal merge.

The regression uses two date prefixes across several partitions and
checks that SELECT DISTINCT returns each date exactly once.

Bug report: https://jira.mariadb.org/browse/MDEV-40820
Yuchen Pei
MDEV-41322 Convert KEY_NOT_FOUND to EOF in "unordered" partition index scans during index_next[_same]/index_prev calls

ha_partition::handle_unordered_scan_next_partition is called in a
variety of accesses, including index_read, index_prev, and index_next.

ha_partition::handle_unordered_next and
ha_partition::handle_unordered_prev are called from index_next and
index_prev accesses. They check for signs of end of
scan (m_part_spec.start_part == NO_CURRENT_PART_ID) and exit early if
that is the case. HA_ERR_END_OF_FILE (EOF) is commonly paired with
assigning NO_CURRENT_PART_ID to m_part_spec.start_part whenever
possible, to signal the end of scan.

The error HA_ERR_KEY_NOT_FOUND means the requested key is not found.
It should not mean the end of scan, when for example
ha_partition::handle_unordered_scan_next_partition is called from
index_read, because a subsequent index_next[_same] / index_prev call
would then incorrectly return immediately from
ha_partition::handle_unordered_next /
ha_partition::handle_unordered_prev. This is why HA_ERR_KEY_NOT_FOUND
is retained and returned in
ha_partition::handle_unordered_scan_next_partition.

But if the call is from index_next[_same] / index_prev,
HA_ERR_KEY_NOT_FOUND should indeed mean end of scan. In this patch, we
ensure this is the case by converting HA_ERR_KEY_NOT_FOUND to EOF in
ha_partition::handle_unordered_next /
ha_partition::handle_unordered_prev.
Yuchen Pei
MDEV-41322 Convert HA_ERR_KEY_NOT_FOUND to HA_ERR_END_OF_FILE in pq-less partition index scan
Prathamesh Hukkeri
MDEV-39307: Fix %f in audit plugin timestamp rendering zero microseconds

The server_audit_timestamp_format %f specifier always rendered zero
microseconds. The server downcast the precise time to seconds before
passing it to audit plugins, and the plugin then re-fetched the time
itself at write time.

Pass the server's high-resolution time through the audit API instead:
- extend mysql_event_general with general_time_microseconds (added in
  MYSQL_AUDIT_INTERFACE_VERSION 0x0304), keeping general_time in
  seconds for backward compatibility
- the server_audit plugin uses event->general_time_microseconds for
  query log entries instead of re-fetching the time at write time

Connection and table events carry no timestamp in the audit API, so the
plugin keeps taking the time at event time for those entries.
bsrikanth-mariadb
MDEV-41340: Push a merged view down like the table it stands for

A multi-table UPDATE/DELETE through a mergeable view wasn't pushed down
into FederatedX at all, even when the view was a transparent stand-in for
one of the engine's own tables. The eligibility check (find_multi_upddel_
handler -> get_fed_table_for_pushdown) ran before JOIN::optimize()'s
DT_MERGE step, so it only ever saw the view's own not-yet-merged
TABLE_LIST placeholder, whose table (if any) isn't FederatedX-backed, and
rejected the whole statement.

Naively fixing that by treating an unmerged view as equivalent to a plain
derived table (defer to its inner SELECT, as already done for anonymous
FROM-list subqueries) is unsafe for a genuinely materialized view: the
statement is reproduced as text via TABLE_LIST::print(), which always
prints a named view by its name, merged or not; a view that needs
materialization (e.g. because of LIMIT) has no such name on the remote
server, so the pushed-down statement referenced a table that only exists
locally, and failed with ER_NO_SUCH_TABLE.

A second, related problem: a merged view's underlying table keeps the
table name/alias it has inside the view's own definition, not the alias
the statement used for the view. If that collides with another reference
to the same table elsewhere in the statement (a self-join through the
view), the printed statement can't tell the two occurrences apart, and
either the remote server rejects it as an ambiguous "not unique
table/alias" statement, or worse, a column silently binds to the wrong
occurrence.

- sql_select.cc: Sql_cmd_dml::execute_inner() now runs the DT_MERGE step
  before asking engines whether they can take over the statement, so a
  merged view's real underlying table(s) are visible to the check instead
  of just the view's own placeholder. Idempotent, so it does not repeat
  the same step inside optimize_inner() further down.
- federatedx_pushdown.cc: get_fed_table_for_pushdown()'s per-table check is
  now the recursive check_fed_table_for_pushdown(), which:
  - skips a merged view's own inert placeholder but recurses into what it
    actually stands for, however DT_MERGE represented it (a spliced
    sibling TABLE_LIST, or a NESTED_JOIN wrapping the placeholder when the
    view is itself one of the statement's targets);
  - rejects a named view that still needs materialization instead of
    deferring to its inner SELECT, closing the ER_NO_SUCH_TABLE gap above;
  - rejects pushdown outright when the same remote table would be printed
    under colliding names, closing the self-join gap above.

Test: federated.federatedx_pushdown_upd_del.
Yuchen Pei
MDEV-41322 Convert KEY_NOT_FOUND to EOF in pq-less partition index scans during index_next/index_prev calls

ha_partition::handle_unordered_scan_next_partition is called in a
variety of accesses, including index_read, index_prev, and index_next.

ha_partition::handle_unordered_next and
ha_partition::handle_unordered_prev are called from index_next and
index_prev accesses. They check for signs of end of
scan (m_part_spec.start_part == NO_CURRENT_PART_ID) and exit early if
that is the case. HA_ERR_END_OF_FILE (EOF) is commonly paired with
assigning NO_CURRENT_PART_ID to m_part_spec.start_part whenever
possible, to signal the end of scan.

The error HA_ERR_KEY_NOT_FOUND means the requested key is not found.
It should not mean the end of scan, when for example
ha_partition::handle_unordered_scan_next_partition is called from
index_read, because a subsequent index_next / index_prev call would
then incorrectly return immediately from
ha_partition::handle_unordered_next /
ha_partition::handle_unordered_prev. This is why HA_ERR_KEY_NOT_FOUND
is retained and returned in
ha_partition::handle_unordered_scan_next_partition.

But if the call is from index_next / index_prev, HA_ERR_KEY_NOT_FOUND
should indeed mean end of scan. In this patch, we ensure this is the
case by converting HA_ERR_KEY_NOT_FOUND to EOF in
ha_partition::handle_unordered_next /
ha_partition::handle_unordered_prev.
Daniel Bartholomew
bump the VERSION
Jan Lindström
MDEV-41075 : Assertion .thd->in_active_multi_stmt in TOI

Clear OPTION_NOT_AUTOCOMMIT/OPTION_BEGIN before opening the table below,
not after. A storage engine may register itself into the "all"
transaction as part of the table open/lock (e.g. InnoDB's
external_lock()), depending on those bits. Clearing them only after the
table is open is too late: the engine has already registered into "all"
using the still-set bits, and since record_gtid is meant to be a standalone
autocommit-style write, nothing will later issue the matching "all"-level
commit to clear that registration and the performance-schema transaction
handle, and it leaks into whatever runs on this THD next.
Jan Lindström
MDEV-40622 : galera.tmp_space_usage fails: Failed to start mysqld.2

Max_tmp_space_used and Tmp_space_used are binlog cache byte counts
that vary by platform. The test compared them against fixed numbers,
so a differing byte count failed the test.

The test now checks that both values stay within
max_tmp_session_space_usage, and that tmp_space_used resets to 0
after change_user. Test only.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Aleksey Midenkov
MDEV-39074 trans_rollback_stmt(THD *): Assertion `! thd->in_sub_stmt' failed.

slave_close_thread_tables() unconditionally called
trans_commit_stmt()/trans_rollback_stmt(), which assert
!thd->in_sub_stmt. A BINLOG statement with malformed base64 payload
executed from an AFTER INSERT trigger hits this: on decode failure,
mysql_client_binlog_statement() sets thd->is_error() and calls
slave_close_thread_tables() from within the trigger's sub-statement.

Guard it with spcont/in_sub_stmt, deferring cleanup to the
enclosing top-level statement. Other callers run only from the
top-level SQL slave applier thread, so this doesn't change their
behavior.

mysql_client_binlog_statement() itself refuses spcont/in_sub_stmt
with ER_SP_BADSTATEMENT before decoding anything. The guard above
is defense-in-depth for its other, top-level-only callers.

1. BINLOG statement executed from a trigger, SF or SP is disabled
by the patch:

Row events (Rows_log_event) open their target table via a
one-shot check that only fires at the top of a fresh statement.
Inside a trigger, the table is never opened this way. It may be
fixed by reusing query_tables, but:

  - find_locked_table() matched only by table name -- could return
    the TABLE instance the enclosing statement was actively writing
    through, not an idle one. Reusing it would require pre-saving its
    state: record[0]/bitmaps/handler/etc. (will deprecate
    ER_CANT_UPDATE_USED_TABLE_IN_SF_OR_TRG)

  - set_stmt_row_injection()/set_time() calls mutated thd->lex, which
    at that point is main_lex -- shared with the enclosing statement,
    not something safe to touch.

Statement events (Query_log_event) run the embedded query via
mysql_parse(): it bundles lex_start(), reset_for_next_command() and
parse_sql() as one unit meant for a genuinely new top-level statement,
not a one-off nested parse.

Calling parse_sql() directly instead avoids that, but then we own
everything mysql_parse() was doing for us: a private LEX and a
private Query_arena (or allocations land on main_lex/whatever arena
is currently active, shared with the enclosing statement), plus
calling mysql_execute_command() ourselves afterwards.

In any case, DML for query_tables cannot be done due to
ER_CANT_UPDATE_USED_TABLE_IN_SF_OR_TRG reasons explained above.

2. PS for BINLOG statement still works and needs a leak fix for
statement events:

A BINLOG statement decoding to a Query_log_event, executed via
PREPARE/EXECUTE, leaked its nested query's allocations onto the PS's
own persistent arena: thd->stmt_arena pointed at it while mysql_parse()
ran the decoded query. Second EXECUTE asserted on ROOT_FLAG_READ_ONLY,
since PROTECT_STATEMENT_MEMROOT marks that arena read-only after a
successful execution.

The fix redirects thd->stmt_arena to thd itself for the duration of
the nested mysql_parse() call, so stmt_arena->is_conventional() reads
true and activate_stmt_arena_if_needed() (called e.g. from
save_leaf_tables()) never redirects allocations to the PS's arena in
the first place.

Harmless for what it's protecting: leaf_tables_exec is normally cached
on the persistent arena so a repeatedly-executed statement's
SELECT_LEX doesn't rebuild it every time, but our SELECT_LEX is torn
down and reparsed fresh (due to mysql_parse() semantics) on every
EXECUTE, so there's nothing to cache here regardless of which arena is
used.
Aleksey Midenkov
MDEV-41050 versioned DELETE via row_end index leaves a row undeleted

1. Aria/MyISAM engines

A system-versioned DELETE on a MyISAM or Aria table indexed by row_end
could leave the last current row undeleted. The delete scans the current
rows via an equality search on row_end = MAX, which is driven by
mi_rnext_same/maria_rnext_same. That function keeps the search's
reference key in lastkey2 and uses the HA_STATE_RNEXT_SAME flag to
remember it has already stored it. This reference-key mechanism is
exactly what lets the scan modify rows it is walking without skipping
them.

Deleting a versioned row is an in-place update of row_end, and the update
reuses lastkey2 as scratch space for the changed key, so it clears
HA_STATE_RNEXT_SAME to request that rnext_same re-store its reference on
the next call. However, TABLE::delete_row wraps the update in
HA_EXTRA_REMEMBER_POS/HA_EXTRA_RESTORE_POS, and RESTORE_POS restored the
whole saved info->update word, resurrecting the HA_STATE_RNEXT_SAME bit
that the update had just cleared.

As a result rnext_same skipped rebuilding its reference key and compared
subsequent keys against the now-overwritten lastkey2, hitting a spurious
end-of-file and terminating the scan one row early, defeating the engine's
own protection against a Halloween-style skip.

Fixed by preserving the current HA_STATE_RNEXT_SAME bit across
RESTORE_POS instead of restoring the stale saved value. See also the HEAP
fix below: same root cause, different per-engine mechanism.

2. HEAP engine

The row loss also reproduces on the MEMORY (HEAP) engine. A system-
versioned DELETE scans the current rows on the row_end index and turns
each delete into an in-place update of row_end, so it modifies the very
index it is walking.

When the changed key is the scanned one (info->lastinx), hp_delete_key()
repositions the cursor but heap_update() leaves info->update untouched, so
HA_STATE_NEXT_FOUND from the preceding heap_rnext() stays set. The next
heap_rnext() then sees current_ptr == 0 with that bit and takes the
"!current_ptr && HA_STATE_NEXT_FOUND" guard as a false end-of-file,
stopping one row early.

Fixed by clearing HA_STATE_NEXT_FOUND when the scanned index key changed.
HA_STATE_AKTIV is kept (unlike heap_delete): the row is updated, not
removed, so a following op must not fail test_active(). The bit is only
set after a heap_rnext(), so a plain single-row UPDATE never reaches this.
See also the Aria/MyISAM fix above: same root cause, different per-engine
mechanism.

3. Why the fix differs per engine, and InnoDB

Aria and HEAP both trace their handler design back to MyISAM, hence the
same root cause (stale scan bookkeeping after an in-place key change) in
all three, fixed at each engine's own bookkeeping spot. InnoDB needs no
fix: its persistent cursor survives concurrent index modification by
design, already covered by this same test under the timestamp
combination (default-storage-engine=innodb), which passes unmodified.
Yuchen Pei
MDEV-41322 Convert KEY_NOT_FOUND to EOF in "unordered" partition index scans during index_next[_same]/index_prev calls

ha_partition::handle_unordered_scan_next_partition is called in a
variety of accesses, including index_read, index_prev, and index_next.

ha_partition::handle_unordered_next and
ha_partition::handle_unordered_prev are called from index_next and
index_prev accesses. They check for signs of end of
scan (m_part_spec.start_part == NO_CURRENT_PART_ID) and exit early if
that is the case. HA_ERR_END_OF_FILE (EOF) is commonly paired with
assigning NO_CURRENT_PART_ID to m_part_spec.start_part whenever
possible, to signal the end of scan.

The error HA_ERR_KEY_NOT_FOUND means the requested key is not found.
It should not mean the end of scan, when for example
ha_partition::handle_unordered_scan_next_partition is called from
index_read, because a subsequent index_next[_same] / index_prev call
would then incorrectly return immediately from
ha_partition::handle_unordered_next /
ha_partition::handle_unordered_prev. This is why HA_ERR_KEY_NOT_FOUND
is retained and returned in
ha_partition::handle_unordered_scan_next_partition.

But if the call is from index_next[_same] / index_prev,
HA_ERR_KEY_NOT_FOUND should indeed mean end of scan. In this patch, we
ensure this is the case by converting HA_ERR_KEY_NOT_FOUND to EOF in
ha_partition::handle_unordered_next /
ha_partition::handle_unordered_prev.
Marko Mäkelä
MDEV-33966: buf_page_make_young() is a contention point

The buf_pool.LRU list needs to reasonably accurately reflect recently
accessed blocks, so that they will not be evicted prematurely.
Because the list is protected by buf_pool.mutex, it is not a good idea
to maintain the position on every page access.

Instead of maintaining the LRU position on each access, we will
decide on each access whether the block qualifies for promotion, and
record the decision in a "promote" flag in the block descriptor. The
buf_flush_page_cleaner() thread as well as some traversal of the
buf_pool.LRU list, all of which hold buf_pool.mutex anyway, will move
the flagged blocks to the "recently used" end of buf_pool.LRU.

A block qualifies only if it is accessed again at least
innodb_old_blocks_time after the reference point, which is its first
access and, after that, its most recent promotion. The age is measured
in whole seconds and checked on the access, not when a sweep reaches
the block, so that a block that a table scan accessed in one burst
will not be promoted, however long it stays in buf_pool.LRU_old. After
a promotion, the age is measured from the time the sweep moved the
block, which can be later than the access that qualified it. The rule
applies at any position in buf_pool.LRU, and the flag stays set until
a sweep moves the block. Thus, a block that qualified outside
buf_pool.LRU_old keeps the flag until a sweep reaches it, which is
usually after it has moved into buf_pool.LRU_old; before, such a
block was moved on any access when freed_page_clock showed that it was
no longer close to the "recently used" end.

buf_page_make_young_if_needed(), buf_page_make_young(),
buf_page_peek_if_too_old(), buf_page_peek_if_young(),
btr_cur_nonleaf_make_young(), buf_page_t::set_accessed(): Replaced by
buf_page_t::touch(), buf_page_t::touch_no_stamp() and
buf_page_t::make_young_if_needed().

buf_pool_t::freed_page_clock, buf_page_t::freed_page_clock: Remove.
This is no longer meaningful in the revised design.
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).FREE_PAGE_CLOCK now
always reports 0. Before the first eviction, buf_LRU_stat_update() now
records statistics intervals, and buf_LRU_evict_from_unzip_LRU() uses
its formula instead of assuming a disk-bound workload.

page_zip_des_t::state: An atomic 16-bit field that will include
the PROMOTE and OLD flags that would more logically belong
to buf_page_t. We maintain them here (along with some
ROW_FORMAT=COMPRESSED specific state that is protected by page
latches) in order to avoid race conditions and unnecessary
memory overhead. m_end moves out to its own field, shrinking the
packed word to 16 bits; n_blobs narrows from 12 bits to 10 to make
room for PROMOTE and OLD, still comfortably above the 744-column
maximum.

The helpers that set or clear one bit of page_zip_des_t::state always
perform the atomic read-modify-write, which is a full barrier on x86.
A caller that can find the bit already in the target state checks it
first: touch_no_stamp() for PROMOTE, buf_LRU_add_block() for OLD, and
page_zip_write_rec() for NONEMPTY. page_zip_des_t::copy_from() and
set_n_blobs_and_empty() skip the fetch_add() when the delta is 0.
page_zip_decompress_clust() counts the BLOB pointers locally and
invokes add_n_blobs() once.

page_zip_des_t::copy_from(): Replaces the copy constructor that
page_zip_copy_recs() invoked. It copies m_end, NONEMPTY and n_blobs
from the source page, but preserves the PROMOTE and OLD flags of the
destination block, because they describe the position of that block
in buf_pool.LRU. Copying OLD from the source would make
buf_pool.LRU_old_len inconsistent with the list. Like
set_n_blobs_and_empty(), it updates the state with a single
fetch_add(), which preserves any concurrent change of PROMOTE or OLD.

buf_page_t::init(): Takes the ROW_FORMAT=COMPRESSED shift size
(ssize) directly and clears the zip descriptor state itself, so
callers (buf_block_t::initialise(), buf_page_init_for_read()) no
longer need a separate page_zip_des_init()/page_zip_set_size() call.

buf_page_t::invalidate(): Replaces buf_block_modify_clock_inc().
Instead of maintaining a 64-bit counter, we will maintain one
comprising 32+16=48 bits, in modify_clock_low,modify_clock_high.
Worst case there will be exactly n<<48 calls to
buf_page_t::invalidate() before some operation such as
btr_pcur_t::restore_position() is executed. Such a count should
be extremely unlikely but not completely impossible. It is worth
noting that the DB_TRX_ID is only 48 bits, and each transaction
start and commit/rollback will consume an identifier.

buf_page_t::modify_clock(): Replaces the read access of
buf_page_t::modify_clock. Assert that the caller is holding
a page latch. Note: because invalidate() and modify_clock() are
protected with buf_pool.mutex or the buf_page_t::lock, there can
be no issue with regard to the atomicity of accessing the 48-bit field.

buf_page_t::access_time: Store the 16-bit buf_pool.access_clock
rather than a 32-bit millisecond ut_time_ms(). It wraps around every
18.2 hours; ages are computed as uint16_t(now - access_time), which
stays correct across that wrap. This avoids any alignment loss: the
adjacent fields modify_clock_low, modify_clock_high, access_time of
32+16+16 bits nicely add up to 64 bits. access_time is stamped on the
first access after the block was initialized, and on each promotion
by make_young_if_needed(), so that the age of a frequently promoted
block stays exact across the wrap.

buf_pool_t::access_clock: uint16_t(my_interval_timer() / 1000000000),
never 0 (see buf_pool_t::now()), refreshed about once per second by
buf_pool_t::refresh_clock() so that page accesses need not read the
system clock.
srv_master_callback() refreshes it. From buf_pool_t::create() until
srv_master_timer is started, and for good when srv_master_timer is
not started (innodb_read_only, innodb_force_recovery>=2,
mariadb-backup), a separate buf_pool_clock_timer refreshes it.

buf_pool_t::access_clock, buf_pool_t::LRU_old_threshold: Located in a
cache line of their own, because they are read on page accesses and
the adjacent buf_pool fields are frequently written.

buf_page_t::touch(): Stamp access_time on the first access, then
invoke touch_no_stamp(). Return whether this was not the first access,
as the result of buf_page_make_young_if_needed() used to be. A
buffer-fix is sufficient, as in MVCC undo page lookups.

buf_page_t::touch_no_stamp(): Set the PROMOTE flag if
innodb_old_blocks_time is 0, or if accessed_at() is at least that old,
at any position of the block in buf_pool.LRU. Because
buf_pool.access_clock advances once per refresh_clock(), about once per
second, a difference of threshold ticks can span less than
innodb_old_blocks_time; the difference must exceed the threshold, so
that a burst of accesses that crosses a refresh is not promoted. Like
btr_cur_nonleaf_make_young(), do not stamp access_time. Once PROMOTE is
set, later accesses only load the state and return.

buf_page_t::make_young_if_needed(): If PROMOTE is set, invoke
buf_page_t::make_young(). This part is inline, so that a sweep pays no
function call for a block that stays in place. It reads the state of
the block once: buf_pool.mutex, which the caller holds, protects OLD,
and only make_young() clears PROMOTE. The block need not be in
buf_pool.LRU_old: while buf_pool.LRU is shorter than
BUF_LRU_OLD_MIN_LEN, buf_pool.LRU_old does not exist, and blocks are
evicted without ever becoming old. The template parameter only_old
leaves a block that is not old in place; buf_do_flush_list_batch()
uses it, because it visits blocks in the order of oldest_modification,
and moving a block that is not old would not save it from eviction.
The template parameter count_not_young selects whether an old block
that was accessed but not flagged is counted in
buf_pool.stat.n_pages_not_made_young; only the eviction sweeps
buf_LRU_free_from_common_LRU_list() and buf_flush_LRU_list_batch() do
this, and a block that a sweep leaves in the list can be counted again
by a later sweep. As before, buf_pool.stat.n_pages_made_young counts
only the moves of old blocks. Because a sweep moves the block, not the
access, Innodb_buffer_pool_pages_made_young can differ from before for
the same workload.

buf_page_t::make_young(): Clear PROMOTE, stamp access_time, and move
the block to the "recently used" end of buf_pool.LRU. The block can be
read-fixed, because buf_pool_t::unzip() copies PROMOTE and OLD from the
compressed-only descriptor and releases buf_pool.mutex during
buf_zip_decompress(). Unlike buf_page_make_young(), we do not skip such
a block: a page access no longer acquires buf_pool.mutex, and the
sweeps hold buf_pool.mutex but no buffer-fix on the block.

buf_LRU_remove_block(): When buf_pool.LRU becomes too short for
buf_pool.LRU_old, clear the OLD flags from buf_pool.LRU_old to the end
of the list, which holds all the old blocks, not the flags of every
block. While buf_pool.LRU_old does not exist, no block in buf_pool.LRU
is old, so on later removals the loop does nothing.

buf_LRU_add_block(): Write the OLD flag only if it changes. In a list
that is too short for buf_pool.LRU_old, assert that buf_pool.LRU_old
is not defined and that the block is not old, instead of writing the
flag.

innodb_old_blocks_time_update(): New sysvar update callback,
replacing a NULL one, that calls buf_pool_t::set_old_threshold_ms(),
so that SET GLOBAL innodb_old_blocks_time also updates the
LRU_old_threshold in seconds that page accesses read. The threshold
is clamped to 65535 seconds, matching the access_time wrap period;
the sysvar itself still accepts up to UINT_MAX32 milliseconds.

buf_pool_t::set_old_threshold_ms(): Round the millisecond threshold
up, not down, to the nearest second. innodb_old_blocks_time is
documented and accepted in milliseconds; flooring instead of ceiling
would make any configured value from 1 to 999 silently behave as 0
(disabled).

buf_page_t::is_accessed(): Renamed accessed_at(), to stop reading as
a boolean. It returns the access_time stamp of the first access or
of the last promotion, in seconds;
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).ACCESS_TIME now reflects
that.

buf_read_ahead_random(): A page now qualifies once accessed_at()
holds, together with either zip.is_promote() or !zip.old().

buf_read_ahead_linear(): Compare access_time stamps as a signed
16-bit difference, not raw unsigned, so the monotonic-access check
stays correct across the access_time wrap. The resolution of the
stamps is 1 second instead of 1 millisecond.

buf_flush_page_cleaner(): Refresh abstime before proceeding to LRU
eviction after an idle period, so that the next my_cond_timedwait()
will not return immediately on a stale deadline.

buf_LRU_scan_and_free_block(): Declare static.

buf_flush_LRU_list_batch(): In a run of blocks that
make_young_if_needed() moves, release and reacquire buf_pool.mutex
after every 512 scanned blocks, except on the first scanned block.
A move costs much less than an eviction, so the stride is longer than
the one of the eviction path.

buf_pool_invalidate(): Define in the same compilation unit with
buf_LRU_scan_and_free_block().

buf_pool.LRU_old_time_threshold: Replaces buf_LRU_old_threshold_ms.

PageConverter::run(): Renamed from fil_iterate(). In debug builds,
initialize and acquire a dummy exclusive latch on the block, so the
assertion in buf_page_t::invalidate() is satisfied; free the latch on
every path, not only on success.

AbstractCallback::m_zip_ssize: Replaces m_zip_size.

page_zip_des_t: Add calc_ssize()/zip_size() helpers, replacing the
zip_size<->ssize conversion duplicated across buf0buf.cc, buf0rea.cc
and row0import.cc.

innodb.buf_lru_scan_resistance: A new big test that checks that pages
of a table that is accessed in one burst are evicted, even if a sweep
reaches them only after innodb_old_blocks_time, that pages accessed
again after that time are promoted, and that pages read while the
buffer pool is being filled are not promoted later.

innodb_zip.n_blobs_700: A new test that stores 700 BLOB pointers on
one ROW_FORMAT=COMPRESSED page, within the 10-bit n_blobs field. The
index_page_splits counter confirms that the clustered index consists
of its root page alone, without depending on which pages remain in the
buffer pool.

Co-Authored-By: Alessandro Vetere <[email protected]>
Dave Gosselin
MDEV-24907:  Empty the source list after merging an Item_equal

With save_merged false, Item_equal::merge_into_list() merges this
object's Item_equal instances into another, linking *this object's
items into the target.  However this allowed callers to free an
item via one list but later read it through the other; a case of
use-after-free.

If save_merged is false then delete the items from the source.
Daniel Bartholomew
bump the VERSION
  • cc-x-codbc-windows: 'dojob pwd if '3.4' == '3.4' ls win32/test SET TEST_DSN=master SET TEST_DRIVER=master SET TEST_PORT=3306 SET TEST_SCHEMA=odbcmaster if '3.4' == '3.4' cd win32/test if '3.4' == '3.4' ctest --output-on-failure' failed -  stdio
Aleksey Midenkov
MDEV-40854 Use of uninitialized table_will_be_deleted in federated engine

MemorySanitizer report:

==851408==WARNING: MemorySanitizer: use-of-uninitialized-value
    #0 ha_federated::end_bulk_insert() storage/federated/ha_federated.cc:2035:30
    #1 mysql_insert(THD*, ...) sql/sql_insert.cc:1258:11
    ...
  Memory was marked as uninitialized
    #0 __msan_allocated_memory
    #1 my_malloc mysys/my_malloc.c:116:7
SUMMARY: MemorySanitizer: use-of-uninitialized-value ... end_bulk_insert()

table_will_be_deleted is a handler member with no in-class
initializer, and the handler object itself is heap-allocated via
my_malloc(), so it starts out as garbage. It was only ever set in
extra(HA_EXTRA_PREPARE_FOR_DROP) and in external_lock().  For a
TEMPORARY table, statement execution can reach
end_bulk_insert()/write_row(), which reads the flag, without
external_lock() having run first, so the read sees uninitialized
memory.

The fix initializes table_will_be_deleted in reset(), which runs
after each statement, so the flag is cleared before the next
statement reaches end_bulk_insert() regardless of locking path.
bsrikanth-mariadb
MDEV-36986: Support tracing array of primitive types

Json_writer had two separate code paths: add_unquoted_str() for
numbers/bool/null, and add_escaped_str() which added the surrounding
quotes itself for strings. Single_line_formatting_helper, which
buffers consecutive array/object elements to decide if they fit on
one line, assumed only strings could ever be buffered and always
wrapped the flushed values in quotes. As a result, arrays of numbers
(e.g. "depends_on_map_bits", "rec_per_key") were incorrectly rendered
with their elements quoted as strings.

Unify both paths into add_escaped_quoted_str(): the caller now hands
over bytes that are already in their final on-the-wire form. String
escaping (json_escape_to_string) writes its own surrounding quotes,
while numbers/bool/null are passed through unquoted, so the one-line
helper just concatenates the buffered payloads on flush instead of
adding quotes itself. A DBUG_ASSERT in add_escaped_quoted_str() now
checks that every payload is already a quoted string or a bare
number/bool/null token, so a caller that violates the contract trips
an assertion in debug builds instead of silently producing invalid
JSON.

Also:
- Fix Json_writer_array::add(ulonglong)/(size_t), which went through
  add_ll() with a cast to longlong and corrupted large unsigned
  values (e.g. ULLONG_MAX); route them through add_ull() instead.
- Fix mysql-test/include/opt_context_schema.inc: "subquery_runs" was
  nested inside the preceding object instead of being a sibling
  member, and "rec_per_key" items are now declared as "number" to
  match the corrected output.
- Update recorded .result files for opt_trace, opt_context_*, and
  subselect_mat_analyze_json to reflect numbers/booleans no longer
  being quoted inside JSON arrays.
- Extend unittest/sql/my_json_writer-t.cc with coverage for arrays
  of primitives: plain integers, mixed types, sizes, multi-line
  arrays, values flushed before a nested object, strings that were
  already escaped by the one-line buffer, and invalid utf8mb4 input
  through add_str().
- Rewrite the "multi-line array of integers" test to actually
  overflow the one-line buffer (9 seven-digit numbers instead of 7),
  since the previous version fit on one line and never exercised the
  element-per-line flush path it claimed to test.
- Drop the now-redundant single-argument add_escaped_quoted_str()
  overload; all call sites already know their length.
- Update json_escape_to_string()'s doc comments (my_json_writer.h,
  sql_json_lib.h) to state that it quotes its output, not just
  escapes it.
Alessandro Vetere
MDEV-39792 InnoDB: ALTER TABLE FORCE triggers assertion "s" in buf_page_get_gen()

When rebuilding a table from ROW_FORMAT=COMPACT or DYNAMIC into
ROW_FORMAT=REDUNDANT, row_merge_buf_add() fetches the full value of an
externally stored (off-page) CHAR column in a multi-byte character set
and pads it to REDUNDANT's fixed local width via
row_merge_buf_redundant_convert(). That helper already dereferences the
BLOB and calls dfield_set_data(), which clears the field's "externally
stored" flag, since the value is now held in full locally.

The "flag externally stored fields" step further down in
row_merge_buf_add() did not know this had happened. It still consulted
the row_ext_t cache built from the original (pre-conversion) record and,
for a column that is not part of the clustered index's unique key,
called dfield_set_ext() again on the very field that had just been
converted, without restoring its data pointer to a valid 20-byte
external reference. row_merge_copy_blobs() would then read the tail of
the padded, space-filled buffer as if it were a BTR_EXTERN_FIELD_REF,
deriving a garbage tablespace id and crashing buf_page_get_gen()'s
fil_space_get() assertion when the alter tried to build the new
clustered index.

Skip the re-flagging step for a field whose "externally stored" flag is
no longer set. row_build() flags every off-page column, and the
row_ext_t cache only holds a subset of those columns, so a field that
is not flagged is either a converted one (already fully local) or one
that the cache does not hold.

With the field no longer re-flagged, the rebuild completes, and the
rebuilt table passes CHECK TABLE with the full column value.

The MDEV-31025 case in innodb.default_row_format_alter failed on
innodb_page_size=4k and 8k: its ROW_FORMAT=REDUNDANT table has eight
utf32 CHAR(255) columns, which CREATE TABLE rejects with
ER_TOO_BIG_ROWSIZE on those page sizes. Derive the number of columns
from the page size, so that the record still exceeds the maximum local
record size and the fixed-length column c is stored externally. The
whole test now passes on every page size.
Kristian Nielsen
MDEV-41264: binlog-in-engine dump thread not stopping at server shutdown

The binlog-in-engine code for binlog dump thread didn't check
should_stop(info, true) to stop on KILL_SERVER when it reaches the end of
the current binlog. This caused the dump threads to hang indefinitely during
SHUTDOWN WAIT FOR ALL SLAVES, and thus the shutdown to hang as well.

Signed-off-by: Kristian Nielsen <[email protected]>
Lawrin Novitsky
The C/C has been updated to the v3.4.11
Marko Mäkelä
MDEV-41010 BACKUP SERVER TO is slower than mariadb-backup

mariadb-backup --backup always uses a dedicated log_copying_thread()
that eagerly copies the log from the server. Let us do the same
in multi-threaded BACKUP SERVER TO, unless innodb_log_file_buffering=OFF
(which prevents arbitrary-size reads from the log file).

backup_sink::id: The thread identifier (0 to CONCURRENT-1)

InnoDB_backup::context::tracked: Log file queue.

innodb_backup_checkpoint_pmem(), innodb_backup_checkpoint():
Enqueue or detach the old log file.

InnoDB_backup::log_track(), Keep copying the log until we
run out of InnoDB data files to copy.
Mohammad Tafzeel Shams
MDEV-37467: InnoDB Instant ALTER TABLE is not crash safe

The hidden metadata record of instant ALTER TABLE was not written
crash-safely, and recovery could fail to roll it back. These are
independent problems.

First, the metadata record may include externally stored BLOB
metadata. The existing BLOB storage path in
btr_store_big_rec_extern_fields() writes the clustered index record
first, with zero BLOB pointers, and only fills in the BLOB pointers
afterwards. If the server is killed after the mini-transaction that
wrote the (incomplete) metadata record was durably committed, but
before the BLOB pointers were written, the table could become
inaccessible on recovery.

Make metadata BLOB storage crash-safe by writing the BLOB pages and
computing their pointers before the metadata record itself is
inserted or updated, so that the record is always written with
complete BLOB pointers. If the server is killed before the metadata
record is written, the already-written BLOB pages are merely
orphaned, which is safe.

Second, trx_undo_report_row_operation() writes the undo log record in
a mini-transaction of its own, which is committed before the
mini-transaction that writes the metadata record. Because
innobase_instant_try() had already updated SYS_COLUMNS and SYS_TABLES
in earlier mini-transactions, a kill in between left a durable undo
log record for the table while the metadata record was unchanged. On
recovery, trx_resurrect_table_locks() would then load the table
definition before the incomplete transaction was rolled back. The
data dictionary described the table as it would be after the
operation, while the metadata record still described it as it was
before, and btr_cur_instant_init() failed on that disagreement.

Write the undo log record of the metadata record in the same
mini-transaction that inserts or updates the record, so that the two
cannot be separated by a crash: until that mini-transaction is
committed, neither of them is durable.

An undo log record is never split between pages. If the DEFAULT
values of the columns being added are large enough that the undo log
record for updating the metadata record would not fit on one page,
innobase_instant_try() would fail. Determine this before the operation
starts, so that it can be performed by another algorithm instead.

Third, the table definition that recovery loads need not correspond to
the metadata record. dict_load_table_one() reads the committed version
of the SYS_TABLES record, and escalates to READ UNCOMMITTED only when
it finds a SYS_COLUMNS record that was written by a transaction that is
still active. The number of SYS_COLUMNS records that dict_load_columns()
reads is derived from SYS_TABLES.N_COLS, which was read from the
committed version. The record of a column that the operation appended is
located after that many records, so it is never read and the operation
goes unnoticed. Only an instant ALTER TABLE that merely appends columns
can escape this way: ADD COLUMN ... FIRST, DROP COLUMN and column
reordering rewrite the SYS_COLUMNS records of already existing columns.

Detect this on the SYS_TABLES record itself, which is located by table
name and therefore does not depend on N_COLS. Every instant ALTER TABLE
that changes the columns updates that record, because
innobase_instant_try() invokes innodb_update_cols().

Fourth, the rollback writes a metadata record that comprises fewer
fields than the table definition describes, because
btr_cur_trim_alter_metadata() shortens it to the number of fields that
it comprised before the operation. That number determines the size of
the null flag bitmap, and hence the position of the array of field
lengths. rec_init_offsets_comp_ordinary() derives it from the record,
while the two functions that write the record derived it from the table
definition and asserted that the two agree.

- btr_store_big_rec_metadata():
  New function to store the off-page columns of a metadata record
  ahead of time. Each BLOB page is allocated and linked in its own
  mini-transaction, and the resulting BLOB pointers are written
  directly into the (heap-resident) index entry. On failure, it frees
  any pages it already allocated and resets the pointers to zero.

- btr_free_big_rec_metadata():
  New helper to free the BLOB pages written by
  btr_store_big_rec_metadata() and reset the entry's BLOB pointers
  to zero, used both on failure inside that function and by its
  callers when the metadata record ends up not being written.

- row_ins_clust_index_entry_low():
  For a metadata entry that needs external storage, convert it to a
  big record and call btr_store_big_rec_metadata() (with
  log_free_check() allowed, since no latches are held yet) before
  inserting the record. On failure, free the metadata BLOBs and
  convert the entry back.

- btr_cur_pessimistic_update():
  When updating a metadata record that requires external storage,
  call btr_store_big_rec_metadata() (without log_free_check(),
  since index and page latches are held) before modifying the record,
  and free the temporary big_rec vector via btr_free_big_rec_metadata()
  or dtuple_big_rec_free() on the various failure/success paths.

- btr_cur_optimistic_insert():
  Remove the special-cased jump to convert_big_rec for metadata
  entries, since their BLOBs are now always stored ahead of time by
  the caller; assert that a metadata entry never needs external
  storage at this point.

- innobase_instant_try():
  Since btr_cur_pessimistic_update() now stores metadata BLOBs
  before updating the record, big_rec is always NULL here; assert
  this instead of calling btr_store_big_rec_extern_fields().

- trx_undo_report_row_operation():
  New parameter caller_mtr. If it is specified, the undo log record
  is written in that mini-transaction, which is never committed or
  restarted here. An undo log page is added within the same
  mini-transaction if the record does not fit on the current one. A
  temporary table never uses the caller's mini-transaction, because
  that would require changing its logging mode. All other callers
  pass NULL and are unaffected.
  Encapsulate the parameters that describe the row change
  (clust_entry, update, cmpl_info, rec, offsets) in the new type
  trx_undo_row_op, which of the fields are set depends on the
  operation, which the type documents.

- btr_cur_ins_lock_and_undo(), btr_cur_upd_lock_and_undo():
  For an instant ALTER TABLE metadata record, pass the
  mini-transaction that is going to insert or modify the record.

- trx_undo_max_rec_size():
  New function to determine the maximum size of an undo log record,
  that is, the space available on an empty undo log page.

- ha_innobase::check_if_supported_inplace_alter():
  Refuse ALGORITHM=INSTANT if the metadata record already exists and
  the undo log record for updating it would exceed
  trx_undo_max_rec_size(). trx_undo_page_report_modify() stores the
  DEFAULT value of each column that is being added in that record in
  full, inline. No such limit applies when the metadata record is
  being inserted, because trx_undo_page_report_insert() writes
  TRX_UNDO_INSERT_METADATA and no field data.

- dict_load_table_one():
  If the SYS_TABLES record was written by a transaction that is still
  active, load the table definition as READ UNCOMMITTED. A
  delete-marked record is excluded, because SYS_TABLES.NAME is the
  clustered index key: RENAME TABLE delete-marks the record of the old
  name, and the definition that corresponds to that name is the one
  that precedes the rename.

- dict_sys_tables_rec_read():
  Report whether the current version of the record was written by a
  transaction that has not been committed. The function already
  determines this in order to decide whether to read an older
  version of the record, and used to discard the answer. Store the
  fields that are read in the new type dict_sys_tables_rec, instead
  of in separate output parameters. A caller that does not want a
  delete-marked record to be reported as not found used to indicate
  that by specifying trx_id as nullptr; that is now the parameter
  skip_deleted.

- dict_load_table_low():
  New parameter uncommitted_rec, which is passed on to
  dict_sys_tables_rec_read().

- rec_get_converted_size_comp_prefix_low(),
  rec_convert_dtuple_to_rec_comp():
  For a record that includes a metadata BLOB, determine the number of
  nullable fields from the tuple, by way of
  dict_index_t::get_n_nullable(), and not from
  dict_index_t::n_nullable. This is what
  rec_init_offsets_comp_ordinary() does, and it is equivalent for a
  tuple that comprises all fields of the index. Relax the assertions
  that required the tuple to comprise all of them.

- Added test in innodb.instant_alter and innodb.instant_alter_crash
  to test normal working of INSTANT ALTER, crash safety and full table.
Kristian Nielsen
mysqltest: Implement --enable_sync_gtid option

The --enable_sync_gtid option switches to use GTID-based
--sync_slave_with_master (eg. using MASTER_GTID_WAIT() instead of
MASTER_POS_WAIT()).

This is useful to run adapt existing test cases for use with
--binlog-storage-engine. But it is also useful in general for replication
using GTID (which is the default).

The option is off by default to not randomly break existing tests.

Signed-off-by: Kristian Nielsen <[email protected]>