Home - Waterfall Grid T-Grid Console Builders Recent Builds Buildslaves Changesources - JSON API - About

Console View


Categories: connectors experimental galera main
Legend:   Passed Failed Warnings Failed Again Running Exception Offline No data

connectors experimental galera main
Dave Gosselin
MDEV-33616:  Allocate the recovery buffer from the heap

recv_sys.tmp_buf comes from malloc() rather than from the large page
allocator.

main.large_pages fails on macOS with "Warning: Memory not freed: 16375"
at shutdown.  recv_sys_t::find_checkpoint() asks for 1048585 bytes,
my_large_malloc() rounds that up to 1064960 and charges the rounded
figure to the server memory accounting, and recv_sys_t::tmp_free()
credits back the 1048585 that was requested.  ut_malloc_dontdump() takes
the size by value, so it has nowhere to report what my_large_malloc()
wrote back.

The rounding happens whenever my_next_large_page_size() finds a reported
large page size at or below the request.  macOS has no huge page
interface for my_get_large_page_sizes() to consult, so its fallback
branch reports the ordinary page size, 16384 on Apple silicon, and the
request is always rounded.  Linux reads the sizes from
/sys/kernel/mm/hugepages, where the smallest entry is usually 2 MiB, and
a 1 MiB request then gets no large page and no rounding.

The buffer has no alignment requirement.  recv_sys_t::parse() copies a
mini-transaction into it when the record is encrypted in the
FORMAT_ENC_11 log, where it is then decrypted in place, or when the
record wraps around the end of the log file, and reads it back as a byte
sequence.  tmp_free() calls std::free() because the member function
recv_sys_t::free() hides the one from <cstdlib>.

log_sys.buf and log_sys.flush_buf keep the large page allocator.  They
round the same way, so a server started with --large-pages
--innodb-log-buffer-size=2101248 still reports 24576 on macOS.  The core
dump exclusion that recv_sys.tmp_buf gives up applies only where
MADV_DONTDUMP exists, so nothing changes on macOS, while a release build
on Linux would now include the buffer in a dump.  tmp_free() overwrites
the redo log records that innodb_encrypt_log decrypted before releasing
the memory, through a volatile function pointer because GCC removes a
plain memset() that is followed by free().

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
ParadoxV5
MDEV-38849 slave_connections_needed_for_purge prevents independent machine from purging binary logs

`@@slave_connections_needed_for_purge`’s default of `1` ensures binary
log availability on replication masters, but is not a sensible default
suitable for all scenarios, especially for long-term slave servers and
standalone (not in a replication setup) servers.
The outcome was that standalone server users were confused why automatic
binlog purging does not work.

This commit changes this default to `0`, which is suitable for both
standalone and (when backed by prompt failure recovery)
replication setups.
`0` also more closely matches the behaviour before MDEV-31404,
which added this variable, out of the box.

This commit also adds a one-time replication warning when registering a
slave, but `@@slave_connections_needed_for_purge` is left unchanged.
Rather than enforcing a defence with an unsensible default, this
reminder will bring awareness of the risk of automatic binlog purging.

This commit also cleans up Galera and MTR workarounds to the
introduction of the `@@slave_connections_needed_for_purge=1` default.
Daniel Black
MDEV-41156: ASAN use-after-poison in JSON_CONTAINS_PATH

The json_depth_array as allocated for 32 elements but accessed
it as though there wasn't a limit.

Called mem_root_dynamic_array_resize_and_get_val to ensure that
the number of elements was presented, and error JE_EOS (out of
space) if allocation exceeded and return its pointer.

With the array resizing called early mem_root_dynamic_array_resize_and_set_val
isn't required.

Removed unused 'value' variable, as value_ptr was always valid.

Reported by: David Korczynski of Ada Logics
DerZc
MDEV-40689 Wrong result: BIT_AND/BIT_OR/BIT_XOR in WINDOW functions over a frame containing NULL

BIT_AND, BIT_OR, and BIT_XOR window functions can return incorrect
values as a sliding frame moves past NULL input rows.

Adding a NULL argument leaves the bit-aggregate state unchanged, but
removing that row unconditionally calls remove_as_window() with
val_int()'s value. The removal path therefore changes state for a row
that never contributed to the aggregate.

Evaluate the departing window argument once and retain its unsigned
value. Call remove_as_window() only when the evaluated argument is
non-NULL. Leave the existing incremental window algorithm and non-window
aggregation path in place.

The regression checks all three bit aggregates over ROWS BETWEEN 1
PRECEDING AND CURRENT ROW with interleaved NULL and non-NULL values,
including removal of a real zero and restoration of the neutral values
after the frame becomes all-NULL.

Bug report: https://jira.mariadb.org/browse/MDEV-40689
Dave Gosselin
MDEV-41211 Remove unused index_read_idx() from both Federated engines

ha_federated::index_read_idx() and ha_federatedx::index_read_idx()
have no callers and do not override a handler method.  Remove them,
and remove the sentence in the comment on each index_read() that says
index_read() calls it.  Comments that describe index_read_idx() now
name index_read() or index_read_idx_map(), whichever does that read.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Dave Gosselin
MDEV-33616:  Take the read lock many times in perfschema.func_mutex

The wait timer can have a granularity coarser than the time an
uncontended read lock is held, so the recorded duration of one lock can
be zero, which reads back as NULL.  This can cause the test to fail with
a false negative.

Take the lock twenty more times at each measurement point, with the
extra statements silent so the recorded result does not change.  The
mutex part of the test already works this way, since one SELECT
produces ten THR_LOCK::mutex events.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Dave Gosselin
MDEV-33616:  Skip the redo log upgrade tests without sparse file support

innodb.log_upgrade and innodb.log_upgrade_101_flags build 8GB redo log
files by seeking past the end of an empty file and writing a single
byte.  That needs a filesystem which leaves the skipped range
unallocated.  HFS on macOS allocates every block of it instead, so the
write fails with ENOSPC and the test reports a perl failure.

include/have_sparse_files.inc probes a directory the caller names,
writing one byte 64MB into an empty file there and comparing the
allocated block count against that offset.
Sergei Petrunia
Add comment about Create_tmp_table::m_group
Alessandro Vetere
MDEV-41201 LRUItr::start(): fix inverted reset condition

The scan hand pointer m_hp must stay in place while it is still
inside the old (cold) sublist of the LRU list, and rewind to the
tail only once it has advanced past the old/young boundary into the
young (hot) sublist. The condition was inverted, so start() rewound
the pointer to the tail on every call instead of only at that
boundary, defeating the intended scan-resumption optimization and
biasing scans toward the tail of the old sublist.

Port of percona/percona-server@dc344fba7e3d272ac18e8d6589b37c678f1e1ad5
by Paweł Olchawa (PS-11446).
Dave Gosselin
MDEV-41303:  rand() in a semi-join subquery is checked on outer rows

Do not merge a subquery into its parent as a semi-join when it has
the UNCACHEABLE_RAND flag, which RAND() and ROWNUM set.  Derived
tables already follow this rule.  ROWNUM sets the same flag, so this
patch replaces the check for ROWNUM with a check for UNCACHEABLE_RAND.

Previously, converting an IN subquery to a semi-join moved its WHERE
into the parent WHERE.  A condition there such as rand(1) < 0.09
doesn't rely on any columns, so it is attached to the last table of
the join order that is outside any materialized semi-join.  With
SJ-Materialization it was checked once for each outer row instead of
once for each row of the subquery, and the query returned a wrong
count.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Dave Gosselin
MDEV-33616:  Make two tests independent of lower_case_table_names

macOS puts the data directory on a case insensitive file system, so
lower_case_table_names is 2 and both tests recorded an answer that only
holds for 0.

period.i_s_notembedded looked up I_S.PERIODS and I_S.KEY_PERIOD_USAGE by
the schema name TEST.  That comparison follows the table name
comparison, so it finds the table under 1 and 2 and finds nothing under
0.  Those four queries move to the new test period.i_s_case_sensitive,
which requires lower_case_table_names=0.  The win rdiff of
period.i_s_notembedded covered the same difference and is no longer
needed.

atomic.drop_db_long_names generated table and view names in upper case
and compared the DROP statements that DDL recovery writes to the binary
log.  Under 2 the names come back from the directory in lower case.
Generating them in lower case to begin with gives the same names on
every setting.  Lower case also changes where the view name sorts
against its table name for the letters after v, which moves one view
between two of the recorded DROP VIEW statements.
Dave Gosselin
MDEV-41211 Federated multi-table DELETE keeps a const table row

A multi-table DELETE on a FEDERATED or FederatedX table went wrong
when a primary key lookup made the target a const table.  The
optimizer reads a const table's row through index_read_idx_map(),
whose default implementation ends the index scan and frees the result
set.  The server asks for the row's position later, during execution,
so the saved position was empty.  FederatedX skipped the row and
FEDERATED crashed in rnd_pos().

Both engines now override index_read_idx_map() so that the lookup
leaves its result set open, as index_read() does.  position() then
records a valid position, and the result set is freed at the end of
the statement.  FEDERATED's reset() now also clears stored_result,
which otherwise pointed at a freed result set and was freed again
when the table was closed.

Tests include a multitable UPDATE with a const target, which crashed
both engines before the fix.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Dave Gosselin
MDEV-33616:  Detect select() on macOS

macOS declares select() in sys/select.h, which the HAVE_SELECT probe did
not include.  clang rejects a call to an undeclared function, so the
probe failed and HAVE_SELECT was left undefined.

my_sleep() then took its last fallback, a busy loop on time() that
rounds the requested interval up to a whole second.  Every sub-second
sleep in the server became a one second spin on a CPU, which is what
made rpl.rpl_perfschema_applier_status_by_worker,
rpl.rpl_shutdown_sighup and rpl.rpl_semi_sync_shutdown_await_ack fail.
Rucha Deodhar
MDEV-41181: ASAN heap-buffer-overflow after SELECT JSON_SCHEMA_VALID
Georg Richter
MDEV-40146 vio_gencert function doesn't set serial number

this apparently breaks RFC 5280, and makes python cryptography module
unhappy.
Dave Gosselin
MDEV-33616:  Match the macOS dlopen error in plugins.multiauth

The client reports why it could not load client_ed25519, and macOS names
every path that dlopen() tried.  Two expressions are added, one for the
chunk that holds the start of that message and one for the chunk that
holds the rest of it.

The line runs to 563 bytes, 52 of prefix and the 511 that the client
error buffer holds, while do_exec() reads the output with fgets() into a
512 byte buffer and runs the replacements on each chunk on its own.  A
long enough vardir therefore splits the line, because the path appears
four times in the dlopen text.  The second chunk is the tail of a path
and carries no colon, where the first chunk keeps the colons of the
mysqltest prefix.  That chunk also holds the only line terminator the
error line gets, so the expression captures the newline and the
replacement puts it back.  A replacement is inserted as written, so a \n
spelled there would reach the output as a backslash and an n.

Both expressions stop at a newline.  reg_replace compiles with
REG_DOTALL, so an unrestricted .* runs past the line terminator whenever
the whole message reaches the replacement in one chunk, and the error
line then joins the line after it.
Vladislav Vaintroub
build mysqlservices without an embedded CRT requirement

mysqlservices only exposes a thin C API, no CRT state crosses it, so
don't force whatever CRT/config built the server onto a plugin linking
it. Without /Zl, a plugin built in a config with no matching installed
mysqlservices variant (CMake silently substitutes one - verified with
a toy project) gets an ignorable but noisy LNK4098 warning.

Assisted-by: Claude:claude-5-sonnet
Dave Gosselin
MDEV-33616:  MTR flag to mark tests as incompatible with macOS

Introduces a new MTR include, not_mac.inc, which when included at the
top of a test, prevents that test from running on macOS.

sys_vars.sysvars_readonly_debug is the first user.  It expects the
server to fault when a read only sysvar is written behind the sysvar
interface.  That protection needs the ro_after_init section, which a
linker script places and ld64 has no option to take, so
HAVE_RO_AFTER_INIT stays undefined on macOS.  Without it no variable is
moved into the read only root either, so neither of the two assignments
is refused.
Dave Gosselin
MDEV-31180:  MyISAMMRG Crash on UPDATE of an updateable VIEW

Attach the children of a MERGE table once per statement, and keep the
value of pos_in_table_list for a MERGE table on subsequent executions
of a prepared statement.
Mohammad Tafzeel Shams
MDEV-41242 : Fix resource leaks on InnoDB/mariabackup error paths found by Infer

Several error-handling paths returned without releasing a resource
already acquired earlier in the function, or checked the wrong handle
entirely, risking use of an unopened handle.

Changes:
- SysTablespace::read_lsn_and_check_flags(): close the datafile handle
  on header-validation failure.
- xb_process_datadir(): check the freshly opened `dir` handle instead
  of the stale `dbdir`, fixing a handle leak and a possible use of an
  unopened directory handle.
- wsrep.cc / xb_load_list_file(): close file handles before die(), and
  null-check fopen() results in wsrep.cc.
- datadir_iter_new(): free datadir_path and destroy the mutex on the
  os_file_opendir() failure path.
Oleksandr Byelkin
MDEV-39993 Use CREATE OR REPLACE for sys schema routines

mariadb-upgrade silently dropped EXECUTE grants on sys schema stored
functions and procedures. The sys schema install scripts reinstalled
every routine with DROP FUNCTION/PROCEDURE IF EXISTS followed by
CREATE. DROP cascades to delete the routine's rows in
mysql.procs_priv, so any EXECUTE grant a DBA had issued on e.g.
sys.table_exists or sys.quote_identifier was lost every time
mariadb-upgrade reinstalled the sys schema, even though the routine
itself came back unchanged.

Fix: replace DROP ... IF EXISTS + CREATE with CREATE OR REPLACE in
all 53 sys_schema function/procedure files and in the two templates
(templates/function.sql, templates/procedure.sql) so future routines
follow the same pattern. CREATE OR REPLACE PROCEDURE/FUNCTION goes
through sp_drop_routine_internal(), which only deletes the
mysql.proc row and never reaches sp_revoke_privileges() (that is
only called from the explicit DROP PROCEDURE/FUNCTION statement),
so mysql.procs_priv is left untouched and existing grants survive.
This mirrors the pattern already used by sys schema views
(CREATE OR REPLACE ... VIEW, since MDEV-9077), which never had this
problem.

Four of the converted files (functions/format_path.sql,
functions/ps_is_account_enabled_57.sql,
procedures/ps_setup_reset_to_default.sql,
procedures/ps_trace_thread_57.sql) are not referenced by
scripts/sys_schema/CMakeLists.txt; they were converted anyway for
consistency and have no behavioural effect.

As a side effect, a pre-existing UDF whose name collides with a sys
routine name is no longer destroyed before the reinstall fails:
DROP FUNCTION IF EXISTS resolved the UDF namespace first, silently
dropping the UDF and then failing on ER_SP_ALREADY_EXISTS anyway;
CREATE OR REPLACE fails immediately on ER_UDF_EXISTS with the UDF
intact.

scripts/maria_add_gis_sp.sql.in and the sys_config triggers were
deliberately left untouched: the GIS procedures are only
re-installed at bootstrap time (mariadb-upgrade instead patches
their definer in place via UPDATE), and triggers carry no
procs_priv rows, so neither is on the code path this bug is about.

Added mysql-test/main/mysql_upgrade_sys_routine_grants.test, which
grants EXECUTE on sys.table_exists and sys.quote_identifier,
overwrites both routine bodies with a marker to prove the upgrade
actually reinstalls them (rather than the sys schema install being
skipped), runs mariadb-upgrade, and checks both that the grants
survived and that the real routine bodies came back.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Dave Gosselin
MDEV-33616:  Routines of a mixed case database are not listed

At lower_case_table_names=2 this returns nothing.

  CREATE DATABASE Db1;
  CREATE FUNCTION Db1.f1(a INT) RETURNS INT RETURN a;
  SELECT ROUTINE_NAME FROM information_schema.ROUTINES
  WHERE ROUTINE_SCHEMA='Db1';

mysql.proc records the function's database as db1, in lower case.
Creating a routine lower-cases its database name whenever
lower_case_table_names is anything but 0, at sql/sp_head.h:121.  The
datadir, SCHEMATA and DATABASE() all keep Db1.

CALL Db1.f1() still works, because calling a routine lower-cases the
database name too and then searches mysql.proc for db1.  The query
above never lower-cases it.  It searches for Db1, and mysql.proc.db
collates utf8mb3_bin, so the comparison runs byte for byte and no row
matches.

At setting 1 the server lower-cases the filter value as well, at
sql/sql_show.cc:4394, and lower-cases every name it stores, so the
query and the table always agree.  Setting 2 lower-cases the routine's
copy and nothing else.

The fix lower-cases the filter value before the search.

Sorting the same query brings the row back.

  SELECT ROUTINE_NAME FROM information_schema.ROUTINES
  WHERE ROUTINE_SCHEMA='Db1' ORDER BY ROUTINE_NAME;

The sort keeps the filter from reaching that search.  The server reads
all of mysql.proc instead, then applies the WHERE to ROUTINE_SCHEMA,
which compares case insensitively.  That shape answered correctly all
along.

The same search fills PARAMETERS and backs SHOW FUNCTION STATUS, SHOW
PROCEDURE STATUS, SHOW PACKAGE STATUS and SHOW PACKAGE BODY STATUS.
Every one returned nothing for Db1.  mariadb-dump lists routines with
SHOW FUNCTION STATUS WHERE Db=..., at client/mysqldump.cc:2859, which
is the main.mysqldump failure.

Setting 0 keeps Db1 and db1 as two databases holding two routines.  A
case sensitive volume confirms both stay distinct before and after this
change.  beb9a5459d4 (MDEV-20609) added the search in 10.11.1.
main.lowercase_routines runs both query shapes.
Brad Smith
crc32c: check elf_aux_info() return value in ppc64 probe

elf_aux_info(3) leaves the output buffer unmodified on failure, so
ignoring the return value could test an uninitialized cpufeatures and
wrongly enable the POWER8 vector-crypto path.

Treat failure as "no features" so the probe falls back to the generic
implementation.
ParadoxV5
Merge branch '11.4' into MDEV-38849
Sergei Golubchik
fix the build for -G "Ninja Multi-Config"
bsrikanth-mariadb
MDEV-36096: Assertion failure in recompute_join_cost_with_limit

An assert was present in method recompute_join_cost_with_limit()
to make sure the recomputed cost is always >= 0.
Although, the assert was correct, it was failing in
CLang compiled versions due to floating point comparison,
when we set  sql_select_limit=1, and
optimizer_join_limit_pref_ratio=1;

In GCC compiled versions, the partial_join_cost was computed to +0.0.
However, in CLang version, the cost turned out to be -0.0.

Changed the assert such that partial_join_cost, would now be checked for
a value >= -DBL_EPSILON. Following it, we set the partial_join_cost to
0, if it has a negative value.
Dave Gosselin
MDEV-28509:  Dereferenced null pointer of type 'struct JOIN_TAB' in add_key_field

setup_group no longer writes Item::marker.  The ONLY_FULL_GROUP_BY
check now tests GROUP BY membership by walking the GROUP BY list.

A query that defines a WINDOW but never refers to it could crash in
add_key_field, for example

  WITH cte AS (SELECT i FROM (SELECT i FROM t1 GROUP BY i) dt
              WINDOW w AS (PARTITION BY i))
  SELECT a.i FROM cte a JOIN cte b ON a.i=b.i WHERE a.i != 5;

A query that defines a WINDOW goes through setup_group, which set
marker to MARKER_UNDEF_POS (-1) on each GROUP BY expression so that
the ONLY_FULL_GROUP_BY check could skip it.  Other code reads marker
as a set of flag bits (-1 sets all bits).
Item_direct_view_ref::grouping_field_transformer_for_where then took
the ref as flagged for substitution and followed a path that ends in
the crash.

The ONLY_FULL_GROUP_BY check was the only reader of that value, so
MARKER_UNDEF_POS is removed.  The necessary check is local to the
setup_group function.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Dave Gosselin
MDEV-33616:  Normalize the strerror text in innodb_fts.index_table

The injected deadlock reaches the client as ER_GET_ERRNO carrying errno
11, and the text comes from my_strerror().  11 is EAGAIN on Linux and
EDEADLK on macOS, so the message reads "Resource temporarily
unavailable" on one and "Resource deadlock avoided" on the other.
Replace the quoted text so the test does not depend on it.
Vladislav Vaintroub
MDEV-39533 Resolve reparse points and enforce MY_NOSYMLINKS on Windows

On Windows, my_realpath() only called GetFullPathName(), does not resolve
symlinks, junctions or mount points, unlike POSIX realpath(). At the
same time, my_open() and my_delete() ignored MY_NOSYMLINKS entirely,
so the symlink-attack protection used for MyISAM/Aria's DATA
DIRECTORY/INDEX DIRECTORY (mi_open()/ma_open(),
my_handler_delete_with_symlink()) was silently absent on Windows.

Fix my_realpath() to actually resolve reparse points: open the
path with CreateFile(), which follows them, and read back the handle's
fully resolved path with GetFinalPathNameByHandle(). As a result,
a missing path now correctly returns 1/ENOENT on Windows too, matching
Linux's realpath()-based behavior, instead of always returning 0.

Make my_open() and my_delete() honor MY_NOSYMLINKS on Windows.
Windows has no per-path-component O_NOFOLLOW equivalent, so instead
this mirrors the HAVE_REALPATH branch of the POSIX
NOSYMLINK_FUNCTION_BODY macro: the caller-supplied name (expected to
already be my_realpath()-resolved) is compared against the actually
opened handle's resolved path, and rejected with ENOTDIR -- the same
errno POSIX uses for this exact "not already canonical" condition --
on a mismatch, whether caused by a TOCTOU symlink swap or by the name
never having been fully resolved to begin with.

As part of enforcing MY_NOSYMLINKS for my_delete(), my_win_unlink()
(formerly in my_delete.c) is rewritten and moved to my_winfile.cc
It opens file once, and verifies no symlinks, before removing it.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Dave Gosselin
MDEV-33616:  Only one of two routines named in a statement is found

With lower_case_table_names 0 the server can have databases Db1 and db1,
each with a function f1.  A single statement naming both databases, like
SELECT Db1.f1(), db1.f1(), reported that db1.f1 does not exist.

The set of routines a statement uses compared its entries without regard
to case.  Only one routine was loaded but the reference to the other
found nothing.  The set now compares its entries exactly, as the routine
cache and the lock manager already do.
Sergei Golubchik
MDEV-40608 MariaDB-devel is incomplete for plugins

This works on Linux and on Windows, with rpm/deb/tar.gz/zip
installations.

For rpm/deb it just works, for tar.gz/zip there is no
standard location, so one needs to configure plugin with

  -DCMAKE_PREFIX_PATH=/pah/to/mariadb/basedir

after that, `cmake --install .` works too, installing in the same
basedir.

`cmake --build . --target package` works, creating rpm/deb/targz/zip
depending on whether it's Linux or Windows and whether -DRPM or -DDEB
was specified.

* create and install mariadb-plugin-config.cmake
* for now it only supports one plugin per project, error out
  if there are many
* deb: move all headers that plugins need to libmariadb-dev,
  together with libmysqlservices.a. At least until we'll
  create mariadb-plugin-dev. Nobody should need huge
  libmariadbd-dev to develop a plugin
* rpm: all in MariaDB-devel already, no changes here
* install wsrep headers too, THD layout depends on WITH_WSREP
* show DBUG_OFF, ENABLED_DEBUG_SYNC, and SAFE_MUTEX to plugins, same
  reason (it doesn't happen automatically as they're not in my_config.h)
* but don't install config.h - high chance of name conflict with other
  projects and it's an exact copy of my_config.h anyway.
* adjust plugin.cmake to work for external plugins
* move server-internal part of it to top-level CMakeLists.txt
* remove double-defined macros from unireg.h (the guard doesn't help
  if unireg.h is included first)
* package plugin metadata as yaml in .tar.gz/.zip

ColumnStore, until fixed, needs a backward-compatibility workaround
Sergei Golubchik
fix errmsg-utf8.txt dependencies for Ninja generator

GenError's custom command must specify headers as OUTPUT,
otherwise ninja cannot deduce that mysqld.cc depends on errmsg-utf8.txt

As a bonus, BYPRODUCTS lists generated files for `ninja clean`
Sergei Golubchik
fix building of external plugins

ARG_DEPENDS might be empty.

tests should be under <pluginname>/ not under whatever build dir happened
to be named.
Oleksandr Byelkin
MDEV-41094 KDF() aliases large iteration/width to weak 32-bit values

KDF() narrowed its iteration-count and key-width arguments without
checking they fit, so values differing by 2^32 aliased to the same
small value and silently derived a much weaker key.

Added checks that the key width fits a 16-bit unsigned value and the
iteration count fits a 32-bit value before narrowing them, rejecting
out-of-range values with an error instead of aliasing.
Dave Gosselin
MDEV-33616:  Exclude innodb_log_file_mmap from sys_vars.sysvars_innodb

Its default value depends on the operating system, ON where the log can
be memory mapped and OFF elsewhere, so the recorded row only holds on
some platforms.  The other variables whose default depends on the
operating system are already excluded the same way.
Dave Gosselin
MDEV-33616:  Widen the block count filter in the buffer pool resize test

The test replaces the number of buffer pool blocks with a fixed value so
that the message is stable.  The pattern only accepted 5.., and macOS
builds without a futex use SUX_LOCK_GENERIC, which enlarges buf_block_t
enough to bring the count down into 4...
Thirunarayanan Balathandayuthapani
MDEV-19574: innodb_stats_method is not honored when innodb_stats_persistent=ON

Problem:
=======
When persistent statistics are enabled (innodb_stats_persistent=ON),
the innodb_stats_method setting is not properly utilized during
statistics calculation.

The statistics collection functions always use a hardcoded default
behavior for NULL value comparison instead of respecting the
configured stats method. This affects the accuracy of
n_diff_key_vals (distinct key count), particularly for
indexes with nullable columns containing NULL values.

Moreover, stat_n_non_null_key_vals[] was never computed for
persistent statistics; it stayed at the 0 that
dict_stats_empty_index() assigns.

With innodb_stats_method=nulls_ignored, innodb_rec_per_key()
therefore always found n_diff <= n_null and reported one record
per key for every index. This impacts the query optimizer,
which makes decisions based on inaccurate cardinality estimates.

Solution:
========
Introduced IndexLevelStats to collect statistics at a specific
B-tree level during index analysis.

Introduced PageStats to collect statistics for leaf page analysis.

Refactored the following functions:
dict_stats_analyze_index_level() to IndexLevelStats::analyze_level()
dict_stats_analyze_index_for_n_prefix() to IndexLevelStats::sample_leaf_pages()
dict_stats_analyze_index_below_cur() to PageStats::scan_below()
dict_stats_scan_page() to PageStats::scan()

The innodb_stats_method value is read once per table in
dict_stats_update_persistent() and passed down, so that all
indexes of a table are analyzed with the same method.

Add the stats method name to stat_description when
innodb_stats_method has a non-default value. The suffix is
dropped when the description is already full.

Added the new stat name n_nonnull_fld01, n_nonnull_fld02, etc.
with a stats description, to indicate how many non-null values
exist for the nth field of the index. This value is retrieved
and stored in the index statistics
in dict_stats_fetch_index_stats_step(). The counts are per
column, not per n-column prefix.

rec_get_n_blob_pages(): Calculate the number of
externally stored pages for a record, using ceiling division
by the usable BLOB page payload (blob_part_size), which differs
between ROW_FORMAT=COMPRESSED (zip_size minus FIL_PAGE_DATA) and
the other formats (srv_page_size minus the BLOB header and the
page trailer). For ROW_FORMAT=COMPRESSED the length in
the field reference is the uncompressed length, so the result is an
upper bound.

When the leaf level is scanned in full, the number of leaf pages that
were scanned is reported as n_leaf_pages for a multi level index.
Before, result.n_leaf_pages was overwritten with
index->stat_n_leaf_pages, which dict_stats_empty_index() had
just set to 1, so every index that took the full scan path
reported n_leaf_pages=1. Single page indexes report 1.
This changes cardinality estimates and
therefore leads to multiple changes in existing test cases.

Non-null values are counted only at the leaf level, since only leaf
pages hold actual records. A full scan of the leaf level counts them
exactly. When the level is sampled, the per column count is derived
from the sampled leaves with the same formula as n_diff:

  n_ordinary_leaf_pages * n_non_null_all_analyzed_pages
                        / n_leaf_pages_to_analyze

This is an estimate for NOT NULL columns as well: the sampled leaves
may hold fewer or more records than the average, and a dive that
stops at a boring page contributes nothing to the sum while still
counting in the divisor.

innodb_rec_per_key(): stat_n_non_null_key_vals[i] holds the
number of records in which the i-th indexed column alone is
not NULL, while what has to be excluded here is the number
of records whose first i+1 columns are all not NULL,
because that is the population which the n-column
prefix statistic stat_n_diff_key_vals[i] has to be
corrected against when innodb_stats_method=nulls_ignored:
with NULLs compared as unequal, every record carrying a NULL
anywhere in the prefix adds a distinct value of its own to n_diff.

PageStats::scan(): n_non_null is accumulated and assigned only
for leaf pages, so that a non-leaf scan cannot leave a node
pointer count behind when scan_below() stops at a boring page
without reaching a leaf.

IndexLevelStats::reset_for_level() also clears n_diff[], and
dict_stats_analyze_index() zero initializes the buffer backing it, so
that a level scan which finds no records (a failed
btr_pcur_open_level(), or a non-leaf page whose first record is not
marked as the leftmost one on the level) leaves n_diff[] at 0 instead
of stale values.

IndexLevelStats::sample_leaf_pages() returns early when the group
boundaries for the prefix are empty, which is the same condition.

IndexLevelStats::analyze_level(): Instead of copying the last record
of the page, retain the latch on the page until the record has been
compared with the first record of the next page

dict_stats_fetch_index_stats_step() no longer resets
stat_n_non_null_key_vals[] while processing an n_diff_pfxNN row:
dict_stats_empty_table() has already cleared the array before the
fetch, and with n_nonnull_fldNN rows now being read too,
that reset would make the result depend on the order in which the
rows arrive.

dict_stats_save(): now static function in dict0stats.cc that
takes the innodb_stats_method value, and is removed from dict0stats.h.
dict_stats_update_persistent() saves the statistics itself, so its
callers no longer have to.

Replaced btr_rec_get_externally_stored_len() with
rec_get_n_blob_pages() in dict0stats.cc.

btr_rec_get_field_ref_offs() and btr_rec_get_field_ref(),
together with the BTR_BLOB_HDR_* macros, were moved from btr0cur.cc
to btr0cur.h so that rec_get_n_blob_pages() can reuse them;

btr_rec_get_field_ref_offs() is now a noexcept function
returning size_t.

Changed stat_n_diff_key_vals and stat_n_non_null_key_vals from
ib_uint64_t* to uint64_t*

len_is_stored(): simplified to a single comparison, which is
equivalent for the unsigned lengths that it is used with.

Removed the unused UNIV_STATS_DEBUG build macro (univ.i) and
turned the DEBUG_PRINTF() helper in dict0stats.cc into an
unconditional no-op
Alessandro Vetere
MDEV-39792 InnoDB: ALTER TABLE FORCE triggers assertion "s" in buf_page_get_gen()

When rebuilding a table from ROW_FORMAT=COMPACT or DYNAMIC into
ROW_FORMAT=REDUNDANT, row_merge_buf_add() fetches the full value of an
externally stored (off-page) CHAR column in a multi-byte character set
and pads it to REDUNDANT's fixed local width via
row_merge_buf_redundant_convert(). That helper already dereferences the
BLOB and calls dfield_set_data(), which clears the field's "externally
stored" flag, since the value is now held in full locally.

The "flag externally stored fields" step further down in
row_merge_buf_add() did not know this had happened. It still consulted
the row_ext_t cache built from the original (pre-conversion) record and,
for a column that is not part of the clustered index's unique key,
called dfield_set_ext() again on the very field that had just been
converted, without restoring its data pointer to a valid 20-byte
external reference. row_merge_copy_blobs() would then read the tail of
the padded, space-filled buffer as if it were a BTR_EXTERN_FIELD_REF,
deriving a garbage tablespace id and crashing buf_page_get_gen()'s
fil_space_get() assertion when the alter tried to build the new
clustered index.

Skip the re-flagging step for a field whose "externally stored" flag is
no longer set. row_build() flags every off-page column, and the
row_ext_t cache only holds a subset of those columns, so a field that
is not flagged is either a converted one (already fully local) or one
that the cache does not hold.

With the field no longer re-flagged, the rebuild completes, and the
rebuilt table passes CHECK TABLE with the full column value.

The MDEV-31025 case in innodb.default_row_format_alter failed on
innodb_page_size=4k and 8k: its ROW_FORMAT=REDUNDANT table has eight
utf32 CHAR(255) columns, which CREATE TABLE rejects with
ER_TOO_BIG_ROWSIZE on those page sizes. Derive the number of columns
from the page size, so that the record still exceeds the maximum local
record size and the fixed-length column c is stored externally. The
whole test now passes on every page size.
ParadoxV5
MDEV-38849 slave_connections_needed_for_purge prevents independent machine from purging binary logs

`@@slave_connections_needed_for_purge`’s default of `1` ensures binary
log availability on replication masters, but is not a sensible default
suitable for all scenarios, especially for long-term slave servers and
standalone (not in a replication setup) servers.
The outcome was that standalone server users were confused why automatic
binlog purging does not work.

This commit changes this default to `0`, which is suitable for both
standalone and (when backed by prompt failure recovery)
replication setups.
`0` also more closely matches the behaviour before MDEV-31404,
which added this variable, out of the box.

This commit also adds a one-time replication warning when registering a
slave, but `@@slave_connections_needed_for_purge` is left unchanged.
Rather than enforcing a defence with an unsensible default, this
reminder will bring awareness of the risk of automatic binlog purging.

This commit also cleans up Galera and MTR workarounds to the
introduction of the `@@slave_connections_needed_for_purge=1` default.
Marko Mäkelä
MDEV-15736 fixup: clang -Wunused-but-set-global