Home - Waterfall Grid T-Grid Console Builders Recent Builds Buildslaves Changesources - JSON API - About

Console View


Categories: connectors experimental galera main
Legend:   Passed Failed Warnings Failed Again Running Exception Offline No data

connectors experimental galera main
Monty
Speed up the query_cache interface

- Do checks inline before calling query_cache.send_result_to_client()

Things tried that failed:
- I tried removing calling lex_start() and reset_for_next_command() if
  query is cached, but this is would cause duplicated, hard to maintain
  code and would require double resets of things if query cache was
  not used. I have documented in sql_parse.cc what would need to be
  reset for this approach to work.
Kristian Nielsen
Validate buffer length in Log_event constructors

This patch goes through all the constructors of Log_event subclasses that
are used to implement read_log_event() and adds error checks on the size of
the event buffer to guard against out-of-buffer reads.

Adds a lot of checks, and fixes some existing checks that were wrong or in
somewhat violation of C++ standard (eg. computing a pointer outside of the
available buffer).

This removes support for reading Table_map_log_event without metadata; such
events were only written by MySQL 5.1 servers going back to 2007.

Fix some missing bits in MAX_LOG_EVENT_HEADER, and update test case
rpl.rpl_packet with a longer query to still exceed max_allowed_packet.

Signed-off-by: Kristian Nielsen <[email protected]>
Daniel Black
MDEV-40949 Missing space after how-to-produce-a-full-stack-trace-for-mariadb link
Rex Johnston
PQ: rebuild a heap temporary table to aria

A worker ships its rows through a temporary table, and when that table filled
the statement failed: nothing moved it to disk, so a result set larger than the
session's temporary-table limit could not be run in the workers at all.

pwt_tmp_table_sink::emit_row() now hands a full container to the server's own
create_internal_tmp_table_from_heap(), which rebuilds it with the on-disk engine
and writes the row that did not fit. Calling that from a worker took three
things, each a thread boundary the conversion never had to cross before.

Temporary-file space. temp_file_size_cb_func() charges whichever thread grows an
on-disk temp table and credits whichever thread shrinks or frees it, and for a
worker's container those are the worker and the manager. The worker now hands
its tmp_space_used over with its other counters -- add_to_status() already
carries the field, thread_func_end() was only zeroing it -- and zeroes its own,
so ~THD finds its books square and the later credit lands where the charge went.

Thread-specific memory. The rebuild allocates into the container, which the
manager frees. The worker measures what the rebuild cost it, takes that off its
own local_memory_used, and hands it to the manager in the sink's cleanup(),
which runs with every worker joined; adding to the manager's counter from the
worker would race with the drain that is running at the time.

The engine's own allocations. Aria allocates MY_THREAD_SPECIFIC unless it is
told the table is not confined to one thread, which is what replication's
HA_OPEN_GLOBAL_TMP_TABLE says. open_tmp_table()'s cross_thread argument now sets
that as well as dropping HA_OPEN_INTERNAL_TABLE, which is what the argument
always meant; only the parallel transport passes it.
create_internal_tmp_table_from_heap() takes it through to the re-open.

That leaves the stack-overflow guard. maria_open() caches a pointer into the
opening thread's my_thread_var for alloc_on_stack(), and a worker's is freed
when its THD is destroyed. handler::rebind_to_thread() is the missing sibling of
rebind_psi(): the same thing TABLE::in_use says to the SQL layer, said to the
engine. Aria overrides it to re-point stack_end_ptr, and the transport calls it
wherever it changes in_use. rebind_psi() cannot serve -- it is about
instrumentation, and is called only where a table leaves the table cache, which
a table handed straight from one thread to another never does.

main.parallel_query_spill covers both sizes and checks the content rather than
only the count: LENGTH(b) is evaluated by the manager, from bytes that came back
through a rebuilt container. main.parallel_query_transport had the old
limitation pinned in it, and is updated.

This commit was prepared with Claude Code: it took the three thread boundaries
one at a time from the assertion each produced, and established that the
manager's read path does not currently dereference stack_end_ptr -- Aria uses it
on key-search, write, delete and blob paths, and the manager scans a keyless
container sequentially -- so the rebind closes a latent hazard rather than a
live one.
Oleg Smirnov
Distribute parallel work more evenly

Reduce SPLIT_THRESHOLD and improve the chunk queue logic
Monty
Speed up the query_cache interface

- Do checks inline before calling query_cache.send_result_to_client()

Things tried that failed:
- I tried removing calling lex_start() and reset_for_next_command() if
  query is cached, but this is would cause duplicated, hard to maintain
  code and would require double resets of things if query cache was
  not used. I have documented in sql_parse.cc what would need to be
  reset for this approach to work.
Rex Johnston
PQ: scan-only prototype, for testing the scan and the transport

Not for the shipping tree. Every hunk is bracketed START PROTOTYPE / END
PROTOTYPE so it can be lifted out again in one pass.

Behind debug_dbug='+d,pwt_scan_only' the workers do nothing but read their
chunk of the driving table and ship the rows. There are no worker JOIN_TABs, no
cloned conditions or select list, and no join in a worker. This thread runs the
plan it would have run serially, and the only difference is that the driving
table's JOIN_TAB reads its rows from the transport rather than from the
handler: do_select() stops diverting, start_scan_only() displaces
read_first_record, and the executor above it is untouched.

Why bother. The gate admits a few per cent of the queries in the test suite,
because almost every refusal it makes is about evaluating cloned Items in
another thread rather than about dividing a scan. Take the evaluation away and
nearly all of them go, so the chunked scan and the transport -- the two layers
underneath -- can be run over the whole suite instead of over that sliver.

How to run it. The scan-only flag on its own does nothing; the workers still
have to be asked for:

  ./mtr --parallel=16 --force --max-test-fail=0 --ignore-parallel-diff \
        --mysqld=--parallel-worker-threads=4 \
        --mysqld=--debug-dbug=+d,pwt_scan_only \
        --suite=main

--ignore-parallel-diff is what makes the run readable: without it a third of
the failures are nothing but the _parallel suffix EXPLAIN adds to the access
type. It does not cover FORMAT=JSON, where the value is quoted, so a handful of
json tests still differ on access_type alone. Never combine it with --record.

For one test, or from inside a test, the flag can be set on a live server:

  SET @sd=@@global.debug_dbug;
  SET GLOBAL debug_dbug='+d,pwt_scan_only';
  ...
  SET GLOBAL debug_dbug=@sd;

Reading the results. main/parallel_query_* fail by construction -- they measure
that the workers ran the join, and here they do not -- as do the environmental
five this tree always fails. What is worth reading is anything with a row
difference. As it stands the whole main suite gives 1358 passes and 39
failures, of which none is a wrong answer: 10 are trace, counter or JSON
EXPLAIN content, 7 are the parallel_query tests, 5 environmental, 2 EXPLAIN
only, 2 an unordered select whose rows arrive in another order, and one an
unordered LIMIT 2 that picks a different pair. No crashes.

What it cannot tell you, which is the more important half. Everything a worker
does with an Item is gone, and that is where the defects have actually been: a
worker THD that did not carry the session's time zone, a condition left only in
a Filesort, a reader that wanted a SQL_SELECT the worker copy did not have, a
materialized subquery re-opened per worker. None of those are reachable here;
most cannot even exist in this shape. Passing in this mode says nothing about
the path that ships. Its value is as a lower layer's test, and as a bisection
tool: a query that answers wrongly in both modes is wrong in the scan or the
transport, and one that answers wrongly only in the full mode is wrong in the
worker join.

Two things the mode needs that the full path gets for free. The displaced
reader is put back by finalize_parallel_workers(), because a JOIN_TAB outlives
one execution -- a prepared statement, a correlated subquery, a routine loop
all run the same plan again -- while the manager does not. And a tab whose
condition was partly pushed into the index is given the whole condition back
for the length of the scan: the pushed half lives in the manager handler's
pushed_idx_cond, which is no longer the handler producing the rows, so it would
otherwise be applied nowhere.

This commit was prepared with Claude Code: it wrote the mode, and found three
defects in it by running the suite -- a flush skipped because the scan loop
ends on HA_ERR_END_OF_FILE, the reader left installed across a re-execution,
and the index-condition pushdown above, which is the same trap the full path
hit from the opposite direction.
Rex Johnston
MDEV-40012 Parallel Query: the manager applies the plan's ORDER BY

A plain ORDER BY was refused, so most sorted queries could not run in the
workers at all. The refusal was right at the time: make_aggr_tables_info()
sorts the driving table's own read for a plan that needs no temporary table,
and the workers take that read over, so the sort lost the scan it was attached
to. This shape has no aggregation table to move it to either -- the terminal
after the driving tab sends straight to the client.

The manager does it instead, over the rows it drains. A sort is the one
post-join step indifferent to the order its input arrives in, which is what
lets a plan that ends in one be handed out in chunks.

It sorts a container of the transport's own layout, not the base table, so the
plan's Filesort cannot be reused: its order names the manager's base-table
fields and the rows are in the container. pwt_row_layout remembers which
shipped column each ORDER element sorts on and builds the equivalent order over
a container's own fields, the same way the group key is rebuilt for the
pre-aggregation containers. The drain collects instead of sending, filesort()
runs, and each row read back takes the path a drained row would have taken:
copy_back_row() into the manager's base-table records, then out. The container
is rebuilt on disk if it fills, which needs none of the cross-thread accounting
a worker's does, this one being written, sorted and read in one thread.

pwt_manager_sort_order() decides whether a plan's sort is ours, and the gate and
the setup ask the same function. Every ORDER BY element has to name a column the
container holds -- a field of a scanned table that the query reads, and so ships.
An expression has no column there, and a sort that returns row ids, unpacks into
other fields or stops early is doing something for the plan beyond ordering; both
are left serial. A plan with an aggregation table is a different shape whose sort
AGGR_OP::end_send() already performs.

pwt_table_conds() had to change with it and the order of its three sources is
load-bearing. add_sorting_to_table() hands tab->select to the Filesort and nulls
tab->select_cond, so a sorted tab keeps its condition only in the filesort -- but
push_index_cond() ran long before and left that SQL_SELECT holding just the
remainder. pre_idx_push_select_cond therefore still wins; the filesort is asked
only when there was no pushdown. Taking them the other way round silently drops
the pushed half, which main.innodb_icp catches as rows the serial plan rejects.

main.parallel_query_sort reads named rows back with query_get_value, which takes
the Nth row as the client received it and so asserts the order rather than the
multiset, ascending and descending, and again once the container has spilled to
disk. main.parallel_query_trace used ORDER BY as its example of a declined query
shape and now uses LIMIT, which is still declined.

A capped SELECT must not run in the workers

With SET sql_select_limit=3, a query the gate accepted returned every row --
500 where the serial plan sends 3. The serial executor enforces a row cap in
end_send(), against unit->lim; the manager's drain sends every row a worker
ships and never consults it. So the gate must refuse any select whose unit
carries a cap, and it was testing the syntax instead of the cap:
limit_params.explicit_limit is only set by a LIMIT clause, while
sql_select_limit is installed by mysql_execute_command() as the default limit
of a top-level SELECT, with explicit_limit still unset.

The gate now tests join->unit->lim.is_unlimited(), which is the very value
end_send() would have enforced. This subsumes the explicit-LIMIT case, and it
is per-unit, so the session cap -- which applies only to the top-level select
-- does not cost a derived table's select its parallel scan.

Found reviewing the ORDER BY commit: its fs->limit != HA_ROWS_MAX check turned
out to be load-bearing for the same reason (the implicit cap becomes a filesort
limit), which raised the question of what protected the unsorted path. Nothing
did.

setup_worker_jointabs() builds each worker tab as a struct copy of the
manager's, and since the manager grew a sort stage the driving tab's copy has
carried the plan's Filesort pointer. Nothing reads it in a worker -- the
driving tab reads through the chunk reader, not join_init_read_record() -- and
worker tabs never run JOIN_TAB::cleanup(), so it is inert today. But
JOIN_TAB::cleanup() deletes the filesort and, through it, the SQL_SELECT
holding the plan's condition, so the copy was one future teardown call away
from a double free. The function's own policy is that manager-owned pointers
are cleared from the copy (select, cache_select, pre_idx_push_select_cond);
filesort and filesort_result now join the list, and pwt_assert_tab_inert()
pins the shape: a filesort is only ever the driving table's, and no sort has
run when workers are set up.

No test: the hazard is latent, reachable only by a teardown path that does not
exist yet. The asserts held over the whole main suite run with
parallel_worker_threads forced on.

This commit was prepared with Claude Code: it built the manager-side sort stage
on the existing container and layout machinery, and found the condition-order
defect above by running the whole main suite with parallel_worker_threads forced
on and separating the row differences from the EXPLAIN ones.

This commit was prepared with Claude Code: it probed the drain path with an
implicit sql_select_limit after noticing that only the sorted shape declined,
and confirmed the serial and parallel row counts diverged.

This commit was prepared with Claude Code, as a hardening it proposed while
reviewing the manager-side ORDER BY commit.
Alexander Barkov
MDEV-39518 Allow prepared statements in stored functions in assignment right hand

Allowing prepared statements in stored functions when
a stored function is used in an assignment right hand.

Both DEFAULT clause of a variable initialization and
the right side of the SET statement are supported:

  CREATE PROCEDURE p1()
  BEGIN
    -- case 1: DEFAULT clause
    DECLARE spvar1 INT DEFAULT f1_with_ps(); -- OK

    -- case 2: SP variable assignment statement
    DECLARE spvar2 INT;
    SET spvar2= f1_with_ps(); -- OK
  END;

- Only assignments to SP variables works for now:
  * SET spvar= func_with_ps(); -- OK
  * SET @uvar= func_with_ps(); -- Error

- Only bare function calls are supported for now. Using a function in
  an expression does not make it PS-safe yet:
    SET v= f1()+0;

- The parser now does not reject PS statements in stored functions.
  PS applicability in stored functions is now detected at run time.
  Note, PS statements in triggers are still prohibited by the parser.

- Functions with PS do not acquire MDL locks on tables, and no MDL is
  taken on the routines themselves either. They work like procedures in
  terms of table opening and routine locking: a concurrent DROP FUNCTION
  can complete while such a function is executing.

- Functions with PS are not replicated as a single `SELECT f1()` call.
  They are replicated per-statement, like procedures.

Helper changes:
- Changing the return result for LEX::sp_variable_declarations_init()
  from void to bool to catch errors in the caller properly.

Misc:
- This patch incorporates fixes for the following bugs found during debugging:
  MDEV-40224,MDEV-40225,MDEV-40226,MDEV-40227,MDEV-40240,
  MDEV-40285,MDEV-40288,MDEV-40315,MDEV-40318,MDEV-40890,MDEV-40900,
  MDEV-40901,MDEV-40913,MDEV-40914,MDEV-41013,MDEV-41015,MDEV-41019

Assisted-by: Claude - reviews and minor clean-ups
Rex Johnston
PQ:  allow workers to pre-aggregate SUM, COUNT, MIN, MAX for the manager
Rex Johnston
Parallel Query: Allow limit where it is the manager limiting (TPC-H Q3)

Claude code figured out it was safe, so lets see.
Rex Johnston
PQ: changes to manage mtr output

1) undo cost discount -> revert mtr plan changes
2) worker thread env fixup -> wrong result during evalution of item
3) add --sorted-result in a few spots
4) add an include/not_parallel_execution.inc, mainly for
    main/innodb_mrr_cpk.test, so far.
5) add in missing parallel_query_index test
Sergei Petrunia
Cleanup in pwt_manager_base, pwt_worker_base.
Kristian Nielsen
MDEV-40760: Stop using common_header_len and post_header_len from FD event

The FORMAT_DESCRIPTION_EVENT contains values common_header_len and
post_header_len array for individual event types, which were intended to
describe the format of following events in the binlog/relaylog. However,
this flexibility has been problematic and never useful:

- The values have never changed in MariaDB since before 10.0, nor in MySQL.

- Having to always have the correct format description event available in
  the code when reading an event causes a lot of complexity (and has
  historically caused numerous bugs).

- The potential for a corrupt or malicious format description event
  requires careful validation that these length values are correct, which
  again has historically seen bugs.

Since the values never change anyway, this patch removed the code that reads
them when reading the format description event, just using the hard-coded
values that are the only valid numbers anyway.

Remove support for ancient row events with post_header_len=6.
Remove support for reading binlogs created by custom MySQL forks with
different common_header_len (ie. MDEV-4645).

Signed-off-by: Kristian Nielsen <[email protected]>
Rex Johnston
PQ: allow secondary index scans
Kristian Nielsen
MDEV-40760: Remove Format_description_log_event::event_type_permutation

The only use of this member was removed back in 2013 (!), but some (now
dead) code remains that accesses the member. This patch removes the member
and the related dead code properly.

Signed-off-by: Kristian Nielsen <[email protected]>
Oleg Smirnov
PQ: add mtr --ignore-parallel-diff

Parallel query execution appends _parallel to the 'type' column of
tabular EXPLAIN/ANALYZE, so an existing test whose plan gets
parallelized fails on that one column alone:

  -17    DERIVED r1      ALL            NULL NULL NULL NULL 2
  +17    DERIVED r1      ALL_parallel    NULL NULL NULL NULL 2

Running the whole suite with --parallel-worker-threads > 0 then buries
the failures that are worth looking at.

Add --ignore-parallel-diff, which makes the result comparison ignore
that suffix: once the two files are known to differ, strip the
suffix from both and compare again.

The text has to be exactly one of ALL, range or index followed
by _parallel, and it has to fill the span between two field
delimiters ('\t', '\n', or end of file).
FORMAT=JSON does not match this criteria, so is not affected.
Rex Johnston
Parallel Query: the worker's JOIN needs its unit

main.parallel_query_join and main.parallel_query_worker_side crashed in a
worker:

  Select_limit_counters::get_select_limit (this=0xa5a5a5a5a5a5ad85)
  join_init_read_record (sql/sql_select.cc:25937)
  sub_select
  pwt_worker::execute_and_handoff

0xa5a5.. is TRASH_ALLOC, so this is memory that was allocated and never
written, not memory that was freed.

setup_worker_join() builds the worker's JOIN by hand, because a worker joins
its own copies of the tables over its own chunk, and it sets the fields the
executor was known to read: the table counts, const_tables, select_lex, and
later join_tab and sum_funcs. join->unit was not among them, and until the
rebase onto current main nothing on the worker's path read it.
join_init_read_record() now does, asking the statement's row limit before it
opens a scan:

  init_table_full_scan_if_needed(tab->table,
                                tab->select ? tab->select->cond : NULL,
                                join->unit->lim.get_select_limit());

so the worker reached it through whatever the mem_root happened to hold.

It is given the manager's own unit rather than a copy. The unit is read and
never written here, and the manager holds it until every worker has been
joined; what it says is also what a worker needs to hear, because the gate
refuses a select that carries a row limit at all.

This commit was prepared with Claude Code: it read the poison pattern as
uninitialised rather than freed, which pointed at the JOIN the workers build
themselves rather than at a lifetime problem, and compared what that function
sets against what the executor now reads.
sjaakola
MDEV-41012 Galera appliers hang with foreign key of types UUID, INET4, INET6

Added second test for testing key collisions from transactions modifying
separate rows
sjaakola
MDEV-41012 Galera appliers hang with foreign key of types UUID, INET4, INET6

wsrep_store_key_val_for_row() built the certification key of a row by
collating the column value whenever the field reports MYSQL_TYPE_STRING
or MYSQL_TYPE_VAR_STRING, taking the collation from Field::charset().

The data types implemented on Field_fbt - UUID, INET6 and INET4
report MYSQL_TYPE_STRING, and their charset() is my_charset_numeric,
which is latin1. Their values are however plain binary and accordingly
get_innobase_type_from_mysql_type() maps them to DATA_FIXBINARY. Their
keys were therefore run through latin1_swedish_ci, which
folds them. That corrupts the key in two ways:

1. A key mismatch for one and the same row. The reference key that
wsrep_rec_get_foreign_key() appends for the parent of a child INSERT is
built from the InnoDB record and is not collated, so it no longer
matched the primary key carried by the parent row's own writeset.
Certification saw no dependency between a child INSERT and a concurrent
parent UPDATE, and two appliers could apply them in parallel causing
a hang or crash.

2. A key collision between distinct rows. The folding is many to one, so
different values collapse onto one key, Certification compares keys byte
for byte, so unrelated rows were treated as the same row. Concurrent
transactions on them certified as a conflict and one was aborted with
ER_LOCK_DEADLOCK.

Fix is for  wsrep_store_key_val_for_row() to skip the collation for
fields that InnoDB stores as binary, using the same condition as
get_innobase_type_from_mysql_type(). This is a no-op for the types that
worked before.
Vladislav Vaintroub
MDEV-38918 Make large pages an explicit per-caller opt-in

my_large_malloc() attempted large pages whenever --large-pages was
enabled, silently rounding the size up and reporting it back via an
in/out parameter. ut_malloc_dontdump() never passed that adjusted
size on to its own callers (the InnoDB redo log buffer and
recv_sys_t::tmp_buf), so freeing later used the original, smaller
size, causing the reported "faux memory leak".

Only the buffer pool and the MyISAM/Aria key caches are documented
to benefit from large pages. Everything else that ended up calling
my_large_malloc() only wanted its "do not dump to core" property and
picked up large pages as an undocumented side effect; those buffers
are also small and sequentially accessed, so they would have gained
little from large pages anyway.

Add MY_TRY_LARGE_PAGES: my_large_malloc() and my_large_virtual_alloc()
now only attempt large pages when a caller passes this flag, instead
of always trying whenever the global option is set. Only the buffer
pool and the key caches pass it. The redo log buffer, tmp_buf, and
row0log.cc's crypt buffers no longer request large pages at all,
which removes the size-rounding bug for them without touching that
code.

my_large_virtual_alloc()'s fallback (no usable large page size) must
also return read-write memory right away, like the Windows large-pages
fallback already does, since my_virtual_mem_commit() is a no-op for
MY_TRY_LARGE_PAGES. my_large_pages_flag is now set once, in
my_init_large_pages(), and never changed thereafter, on any platform.

Both my_virtual_mem_commit() and my_virtual_mem_decommit() are now
no-ops, aside from accounting, whenever large pages are requested.

Also fix a broken mtr suppression regex in main.large_pages that
would fail the test on Windows.
Jan Lindström
MDEV-41017 : Galera test failure galera.galera_toi_alter_auto_increment

New warning was added in MDEV-33660 commit bd127faa. Re-record
test result file.
Rex Johnston
Parallel Query: AVG, and an aggregate inside an expression

TPC-H Q1 now runs in the workers. Two things were stopping it, and both are
here.

AVG
---
An average does not compose out of averages, so what a worker has for a group
is two numbers: the sum of its rows and how many there were. Nothing had to be
invented to carry them. Item_sum_avg::create_tmp_field() already packs a sum
and a count into one field, because that is how the serial grouped plan
accumulates an average, and a worker's grouped container was already
maintaining both.

What was missing was the merge. The manager primes its own aggregate with a
worker's partial through direct_add(), and Item_sum_sum::direct_add() carries
only the sum -- Item_sum_avg::count was left untouched, so every partial
counted as one row. direct_add() now takes the count alongside the sum, and
reset_field()/update_field() add it to the field's count instead of the 1 an
ordinary row contributes, in the same shape Item_sum_sum's already had.

Nothing changes for a serial plan: direct_added is false there, and the new
branches are not taken.

An aggregate inside an expression
---------------------------------
SUM(x)+1 was refused, and so was SUM(x)/COUNT(x) -- the obvious way to write an
average by hand. A select list item that is not itself an aggregate is
evaluated once per group from whatever row the terminal is holding, so a field
in it that GROUP BY does not name takes an arbitrary value; that is worth
refusing. But the check walked into the aggregates as well, found x under
SUM(x), and refused a query that was never in doubt.

It now steps over Item_sum subtrees, using with_sum_func() to answer an
aggregate-free subtree wholesale with the walk that was there before, and
decomposing only the node kinds that can hold an aggregate. Anything else
carrying one is refused rather than guessed at: being wrong in the permissive
direction here is wrong results. SUM(x)+x is still correctly refused, because
the bare x is walked and is not a group column.

The shipped row's shape
-----------------------
With more than one AVG the second onward came back wrong, which was not
arithmetic. Item_sum_avg is the one aggregate whose create_tmp_field() differs
between a keyed and an unkeyed temporary table -- it packs the count beside the
sum only for the keyed form -- and the two ends of the transport were built
differently:

  recv, group_container  make_container(..., group)  16-byte {sum, count}
  exec.result (shipped)  make_container(...)          plain 8-byte double

flush_groups() copies a group container record into the shipping container by
reclength, so it wrote past a narrower record and the manager read back a wider
one, drifting by one AVG's worth of bytes at a time. For SUM, COUNT, MIN and
MAX the two widths are identical, which is why this sat there unnoticed. The
shipping container is now built with the same key; a worker ships one row per
group, so there is nothing for the key to collapse.

main.parallel_query_aggregate gains an AVG section -- serial against parallel
over a DECIMAL average, and three AVGs at once, which is the case that exposed
the shape bug -- and AVG leaves its "what is not accepted" list. STDDEV,
BIT_XOR and COUNT(DISTINCT) stay there.

This commit was prepared with Claude Code: it bisected TPC-H Q1 down to the two
refusals, found the container shape mismatch from the pattern that only the
last AVG was wrong, and fixed a null Item passed to VDec that the aggregate
test caught as a crash.
Kristian Nielsen
MDEV-40760: Remove unused Format_description_log_event * arguments

Follow-up patch, remove a lot of Format_description_log_event * function
arguments that are no longer needed.

Signed-off-by: Kristian Nielsen <[email protected]>
Sergei Petrunia
Continue pwt_manager_base cleanup: introduce locked__process_manager_wakeup().

Note this one:
  //  TODO: the following was done when not holding LOCK_data. Does it
  // matter?
Rex Johnston
Parallel Query: give the worker back optimizer_replay_context too

main.parallel_query_clone aborted in safemalloc while a worker's THD was being
destroyed:

  free_memory (mysys/safemalloc.c:275)  DBUG_ASSERT(irem->marker == MAGICSTART)
  my_free
  plugin_thdvar_cleanup (sql/sql_plugin.cc:3429)
  THD::free_connection -> ~THD -> destroy_background_thd
  pwt_thread::init_and_run_thread_func

A cleared marker is a second free of something already freed.

A worker THD takes the session's variables wholesale, so that an item
evaluated in a worker sees the session's time zone and the rest, and then puts
back the members a THD allocates and frees for itself. That list has to be
plugin_thdvar_cleanup()'s list, and it had fallen one behind it: the rebase
onto current main added

  my_free(thd->variables.optimizer_replay_context);

alongside the redirect_url and default_master_connection frees. It is a
SESSION_ONLY Sys_var_charptr, so it is allocated per THD -- but it was not
being restored, so every worker carried away the manager's pointer, the first
worker to exit freed it, and the next free of the same pointer found the marker
gone.

Restored with the others, and the invariant that was only implied is now
written down next to the list: whatever plugin_thdvar_cleanup() frees, this has
to give back. The failure is a long way from either end, so the next member
added there will not announce itself.

This commit was prepared with Claude Code: it read the abort as a double free
rather than a cross-thread free, and diffed what plugin_thdvar_cleanup() frees
against what the worker restores to find the member the rebase had added.
Sergei Petrunia
Add comments

Factor out common code into get_mvi_index()

Factor out common code into Item_func_json_contains::get_mvi_access()

Item_func_json_contains::mvi_analyze() and ::create_ft_for_mvi() were
near-identical: both checked the arguments, looked up the matching MVI,
parsed the constant second argument and ran the same scan loop calling
encode_mvi_key(). They differed only in what they did with each encoded
key.

Move all of that into get_mvi_access(), which returns an Mvi_access, and
give Mvi_access two methods:

- add_key(), to collect one encoded element key,
- create_ft_item(), to build the

    MATCH vcol AGAINST ('+encoded_foo +encoded_bar ...' IN BOOLEAN MODE)

  item. It honors Mvi_access::conjunctive, so JSON_OVERLAPS will get the
  OR form for free.

mvi_analyze() and create_ft_for_mvi() are now thin wrappers around
get_mvi_access().

This also fixes a memory leak: the encoded keys were copied with
String::copy(), giving each String in Mvi_access::encoded a heap buffer
that is never freed (the Strings live on the MEM_ROOT, so their
destructors never run). Copy the keys onto the MEM_ROOT instead.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

Make the MVI scan a real access method: QUICK_MVI_SELECT

JSON_CONTAINS() over a multi-valued index used to be optimized by
rewriting the WHERE clause: setup_mvi_for_join() injected a synthetic

  MATCH vcol AGAINST ('+k1 +k2' IN BOOLEAN MODE)

into join->conds and into select_lex->ftfunc_list, and the normal
fulltext machinery then picked it up as JT_FT access. The injected item
showed up in the plan and in the condition even though the user never
wrote it, and because it became ordinary ref access the scan was never
costed against the alternatives - it won by being in the WHERE clause.

Introduce QUICK_MVI_SELECT (QS_TYPE_MVI), a QUICK_SELECT_I that drives
the fulltext index directly through the handler API. Unlike FT_SELECT
there is no Item_func_match to have created the FT_INFO, so the quick
select creates it in reset() with ft_init_ext() and frees it with
close_search() in its destructor. Mvi_access::create_ft_item() is
replaced by build_ft_query(), which builds just the query string.

The analysis in setup_mvi_quick() is now kept: Mvi_context moves to the
header, is allocated on the mem_root and stored as JOIN::mvi_ctx, where
JOIN::get_mvi_access_for_table() looks it up. get_quick_record_count()
builds the quick select before test_quick_select() and keeps whichever
of the two is cheaper; test_quick_select() itself is untouched, so the
MVI quick is held in a local across the call (it deletes select->quick
on entry). The same save/compare is done around the second
test_quick_select() call in make_join_select(), which a LIMIT can reach.

A fulltext key never gets a bit in const_keys or keys, so mark the MVI
key of every table that has an access: the const_keys bit is what lets
the range analysis run for that table at all, the keys bit puts the
index into EXPLAIN's possible_keys.

Collect the accesses from the top-level AND-parts of the WHERE clause
only, instead of walking the whole condition. An MVI scan reads just the
rows the index matches, so it is only valid for a predicate that must
hold for every row of the result: for

  json_contains(j1->'$.tags','"a"') OR json_contains(j2->'$.tags','"a"')

scanning either index would drop the rows that only match the other
branch. The deleted add_ft_for_mvi() refused COND_OR_FUNC for the same
reason; walking the condition tree lost that, which only became visible
once the accesses were actually used.

Costs are placeholders (records=10, read_time=0.001) until the engine
can estimate a fulltext search. Note that while there is no estimate,
an MVI access is also taken when test_quick_select() produced no quick
select at all, without comparing it to the cost of a table scan.

TODO: This doesn't handle UPDATE/DELETE!  Should it be put into
  check_quick() call?
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Sergei Petrunia
Rename pwt_worker_base2->pwt_worker_base, pwt_manager_base2->pwt_manager_base
Rex Johnston
Parallel Query: count a worker in under the lock that counts it out

A query could hang for good, with the manager waiting on COND_data_avail for a
worker that had already finished and gone.

active_workers is decremented in report_worker_final_state() under LOCK_data,
and read in locked__process_manager_wakeup() with LOCK_data held, but
register_worker() was incrementing it under no lock at all. The two do run at
the same time: init_parallel_workers() builds the team one worker at a time and
starts each thread as it creates it, so worker 1 can be counting itself out
while worker 2 is being counted in.

  manager thread                      worker 1
  ------------------------------      ---------------------------------
  new pwt_worker  -> reads 2
                                      thread_func_end()
                                        LOCK_data, reads 2, writes 1
  writes 3

The decrement is lost and the count stands one too high. From then on
claim_next_result() waits for a worker that will never publish a result and
never signal again, because every worker there was has already exited.

Taking LOCK_data around the increment is the whole fix.

It wants a small table and enough load: a table the engine divides into one
chunk leaves the other workers with nothing to scan, so they reach
thread_func_end() almost at once, in the window while the rest of the team is
still being built. main.function_defaults_innodb is such a test, and this
reproduced it within six attempts:

  ./mtr --parallel=16 --mem --force --ps \
        --mysqld=--parallel-worker-threads=4 \
        main.function_defaults_innodb main.function_defaults_innodb \
        main.function_defaults_innodb main.function_defaults_innodb \
        main.function_defaults_innodb main.function_defaults_innodb

It now runs clean twelve times over. No test is added: a lost update needs a
race to be lost, and a test that only sometimes fails is worse than none.

This commit was prepared with Claude Code: it reproduced the hang, read the
manager's bookkeeping out of the stopped server with gdb -- active_workers 1
against four workers, all four with a destroyed THD and the statistics that
only thread_func_end() writes, so all four had decremented -- and reasoned back
from four decrements and a count of one to a lost increment.
PranavKTiwari
MDEV-25515 User Account Host Names using CIDR notation
CREATE USER u@'192.168.0.0/24';  -- same as '192.168.0.0/255.255.255.0'
CIDR is an input spelling only, normalised to the netmask form while parsing.
A malformed mask is now rejected instead of stored, and one already in
the privilege tables is skipped at load with a warning instead of being honoured.
Nothing rejected these before, and they are not merely
useless: a non-contiguous mask such as 10.0.0.0/255.0.255.0 matches the
scattered set 10.*.0.*.  Skipping rather than erroring keeps such a row
droppable and renamable.
New error ER_INVALID_HOST_NETMASK.
Dmitry Shulga
MDEV-40076: Follow-up commit to fix compiler error on Win64
Sergei Petrunia
Rename pwt_manager_base->pwt_thread_manager, pwt_worker_base->pwt_thread
Rex Johnston
Parallel Query: the gate says why it turned a query away

"Why did this not run in parallel" was a question the reader had to answer
themselves, from a plan, in a debugger, some way from the query that provoked
it. It is now a question the server answers:

  SELECT JSON_EXTRACT(trace, '$**.parallel_scan_declined_because')
    FROM information_schema.optimizer_trace;

  ["grouped: the aggregation table is keyed by a unique constraint rather than
    by the group -- create_tmp_table() does that for a GROUP BY element too
    wide for a key part, and a hash of the row is not something a group can be
    looked up in",
  "the plan needs a temporary table and the workers cannot pre-aggregate into
    it"]

The first entry is the check that objected; the rest is that refusal working
its way back up. Every one of the gate's refusals now goes through
pwt_decline(), which keeps the DBUG_PRINT and adds the reason to the trace.
Sixteen of them had neither before -- a bare "return false" -- and those were
in the two functions a grouped query spends most of its time in.

Three things this needed beyond saying the words.

The gate is asked twice. Once from make_join_readinfo() while the plan is still
being built, and again from parallel_join_check() when it is finished. The
early call runs before make_aggr_tables_info() has built the aggregation table,
so it reported "the terminal is not end_update()" for every grouped query --
an answer about a plan that no longer exists by the time the query runs, and a
worse place to start than no answer at all. A trace flag threaded through the
gate means only the decision that stands speaks. Threaded rather than kept in a
static, because several connections optimize at once.

Which check fires first decides what you are told. Two reasons were being lost
to a coarser test reached earlier: ordered_index_usage now precedes the
access-method test, since "an index supplies the ORDER BY order" says more than
"this is not a scan the engine divides", and both refuse the same plans. And
the duplicated access-method test in table_can_be_parallel_scanned() is gone,
so is_parallel_scan_applicable() owns that question alone and can tell its
several answers apart.

A reason has to name the cause, not the symptom. A GROUP BY on a column too
wide for a key part is refused because create_tmp_table() keys the container by
a unique constraint, which makes the terminal end_unique_update(), which is why
pwt_plan_group_key() finds nothing -- and "the terminal is not end_update()" is
a true statement that sends the reader nowhere. The check that would have named
the width is never even reached. So when there is no group key, this asks
make_aggr_tables_info()'s own condition of the same table it asked it of, and
says which of the two it is.

main.parallel_query_why pins one refusal per shape that can reach it, so a
reason cannot quietly go missing or attach itself to the wrong query.

This commit was prepared with Claude Code, after a report that finding why a
wide GROUP BY column would not run in parallel took an afternoon in gdb. It
counted the refusal points first -- 22 of 38 traced, 16 silent, the one being
hunted among the silent -- and worked from there.
Rex Johnston
PQ  remove accumulate_group(), use end_update() from sql_select.cc

Now we actually *have* to solve the aria table conversion, not
sidestep it.
Rex Johnston
PQ: tidy up and many constraint updates
sjaakola
MDEV-41012 Galera appliers hang with foreign key of types UUID, INET4, INET6

Added a deterministic test for reproducing the issue
Monty
MDEV-39949 Add per-table "SQL_CACHE" option

Add SQL_CACHE=0|1 option for tables to disable/enable query caches if
table is part of the query.
SQL_CACHE option can be disable for a table by using SQL_CACHE=DEFAULT

Also provide a workaround for the problem described at:
MDEV-6631 "No cache directive(SQL_CACHE) query still wait for
query cache lock" on query_cache_type=DEMAND mode"
by adding DEMAND_STRICT mode.
The new mode is needed as modifying how DEMAND works would made the
change backward incompatible.

query_cache_type has now following options (0-2 are same as before)
OFF    (0) = Query cache off
ON    (1) = Query cache on, all queries are cached except queries marked
            with SQL_NO_CACHE
DEMAND (2) = Query cache on, only queries marked with SQL_CACHE are cached
DEMAND_STRICT
      (3) = Query cache on, only queries starting with SELECT SQL_CACHE
            are cached. This avoids a query cache lock+lookup for every
            query.
TABLES (4) = Query cache on, only cache queries marked with SQL_CACHE or
            where all used tables are using the create option SQL_CACHE=1

Note that if any of the tables used in the query has the create option
SQL_CACHE=0 or is a system table, temporary table or other non
cacheable table, then the query will not be cached.

SHOW CREATE TABLE prints SQL_CACHE=0|1 when the option is set. For tables
where the option can never apply (temporary tables, tables in the mysql
database, and engines that do not support the query cache) it is printed
as /* SQL_CACHE=1 */.

co-author: Seunguck Lee
- Took some changes related to MDEV-6631 from his patch.
Alessandro Vetere
MDEV-38056 Assertion 'bpage->state() >= buf_page_t::UNFIXED' in buf_page_get_zip()

Three callers reconstruct old versions of a clustered index record, derive
secondary index entries from them, and dereference the BLOB pointers of
externally stored columns: row_vers_impl_x_locked_low() for the implicit lock
check, row_undo_mod_sec_is_unsafe() for rollback, and row_check_index() for
CHECK TABLE ... EXTENDED. purge_sys.view is what keeps those pages allocated,
but purge_sys_t::view_guard froze it only for the duration of
trx_undo_prev_version_build(), and the dereference happens after. On
ROW_FORMAT=COMPRESSED this trips the assertion above; in a release build the
freed page is read anyway, which is silent corruption.

All three now hold purge_sys.latch across the dereference, and only where the
version has externally stored columns, since row_build() dereferences nothing
otherwise. In row_check_index() the freeze spans the whole comparison, which
fetches one field at a time and may evaluate a virtual column expression; for
the latest version it is also confined to a delete-marked record, the only
kind that need not own what it points at. Freezing that late means the view
may have advanced since the walk decided a version was reachable, so each
caller re-establishes that under the freeze; trx_undo_prev_version_build()
states the condition and why testing the oldest writer applied suffices.

The first two callers can treat that test as an invariant and end the walk
where it fails. row_check_index() cannot, because it decides reachability from
the lagging purge_sys.end_view on purpose, and that lag is how it finds orphan
secondary index records: where the test fails it stops and reports nothing
rather than treating the record as an orphan, and no report is lost for good,
end_view eventually reaching the same point. Its two purge_sys.is_purgeable()
tests now read the frozen view wherever a fetch follows, which makes them
atomic with the fetch they guard.

trx_undo_prev_version_build(): remove the gate that was meant to stop CHECK
TABLE ... EXTENDED from fetching BLOBs it may no longer own, and
view_guard::is_extended() with it, which no guard mode could satisfy. That
decision belongs to the caller, the only one that knows whether it will
dereference anything.

row_log_table_get_pk(): document why the online ALTER path may dereference
without a freeze.

Debug-only keywords. purge_hold_cleanup parks a purge batch between its last
purged record and purge_sys_t::batch_cleanup(), the window in which a reader
that goes by purge_sys.end_view can still reach history the batch has removed;
a batch opens and closes it without ever returning to the test.
purge_no_blob_freeze sends a version that does have externally stored columns
down the path one without any takes, which is what all four call sites did
before this change, and reproduces the assertion above.
row_vers_impl_x_locked_purgeable, row_undo_mod_sec_is_unsafe_purgeable and
row_check_index_purgeable force the re-validation to fail, reaching exits that
purge_sys.view advancing mid-walk otherwise produces. The first reports no
implicit lock for a row that a live transaction still holds, so a test may
only check that nothing breaks; the other two make the server more cautious.

Tests. old_blob and old_blob_updel cover the implicit lock check,
old_blob_rollback the rollback, old_blob_check CHECK TABLE ... EXTENDED. Each
parks a walk at a dereference, makes the BLOB freeable, and asserts that the
counter of purged update records, which is what would free it, stays at zero
while parked and advances once the walk is over. old_blob_updel covers an undo
log record that stores only the 20-byte reference and needs purge_hold_cleanup
to reach it; old_blob_rollback parks at a reference that the version merely
inherited. All four fail with the original assertion under
debug_dbug=+d,purge_no_blob_freeze, and skip above a 16k page size, which
ROW_FORMAT=COMPRESSED requires. old_blob_purgeable drives the three
re-validation exits, needs no synchronisation, and runs at every page size.
Monty
MDEV-39949 Add per-table "SQL_CACHE" option

Add SQL_CACHE=0|1 option for tables to disable/enable query caches if
table is part of the query.
SQL_CACHE option can be disable for a table by using SQL_CACHE=DEFAULT

Also provide a workaround for the problem described at:
MDEV-6631 "No cache directive(SQL_CACHE) query still wait for
query cache lock" on query_cache_type=DEMAND mode"
by adding DEMAND_STRICT mode.
The new mode is needed as modifying how DEMAND works would made the
change backward incompatible.

Query_cache_type has now following options (0-2 are same as before)
OFF    (0) = Query cache off
ON    (1) = Query cache on, all queries are cached except queries marked
            with SQL_NO_CACHE
DEMAND (2) = Query cache on, only queries marked with SQL_CACHE are cached
DEMAND_STRICT
      (3) = Query cache on, only queries starting with SELECT SQL_CACHE
            are cached. This avoids a query cache lock+lookup for every
            query.
TABLES (4) = Query cache on, only cache queries marked with SQL_CACHE or
            where all used tables are using the create option SQL_CACHE=1

Note that if any of the tables used in the query has the create option
SQL_CACHE=0 or is a system table, temporary table or other non
cacheable table, then the query will not be cached.

SHOW CREATE TABLE prints SQL_CACHE=0|1 when the option is set. For tables
where the option can never apply (temporary tables, tables in the mysql
database, and engines that do not support the query cache) it is printed
as /* SQL_CACHE=1 */.

co-author: Seunguck Lee
- Took some changes related to MDEV-6631 from his patch.