Branch data Line data Source code
1 : : /*-------------------------------------------------------------------------
2 : : * worker.c
3 : : * PostgreSQL logical replication worker (apply)
4 : : *
5 : : * Copyright (c) 2016-2026, PostgreSQL Global Development Group
6 : : *
7 : : * IDENTIFICATION
8 : : * src/backend/replication/logical/worker.c
9 : : *
10 : : * NOTES
11 : : * This file contains the worker which applies logical changes as they come
12 : : * from remote logical replication stream.
13 : : *
14 : : * The main worker (apply) is started by logical replication worker
15 : : * launcher for every enabled subscription in a database. It uses
16 : : * walsender protocol to communicate with publisher.
17 : : *
18 : : * This module includes server facing code and shares libpqwalreceiver
19 : : * module with walreceiver for providing the libpq specific functionality.
20 : : *
21 : : *
22 : : * STREAMED TRANSACTIONS
23 : : * ---------------------
24 : : * Streamed transactions (large transactions exceeding a memory limit on the
25 : : * upstream) are applied using one of two approaches:
26 : : *
27 : : * 1) Write to temporary files and apply when the final commit arrives
28 : : *
29 : : * This approach is used when the user has set the subscription's streaming
30 : : * option as on.
31 : : *
32 : : * Unlike the regular (non-streamed) case, handling streamed transactions has
33 : : * to handle aborts of both the toplevel transaction and subtransactions. This
34 : : * is achieved by tracking offsets for subtransactions, which is then used
35 : : * to truncate the file with serialized changes.
36 : : *
37 : : * The files are placed in tmp file directory by default, and the filenames
38 : : * include both the XID of the toplevel transaction and OID of the
39 : : * subscription. This is necessary so that different workers processing a
40 : : * remote transaction with the same XID doesn't interfere.
41 : : *
42 : : * We use BufFiles instead of using normal temporary files because (a) the
43 : : * BufFile infrastructure supports temporary files that exceed the OS file size
44 : : * limit, (b) provides a way for automatic clean up on the error and (c) provides
45 : : * a way to survive these files across local transactions and allow to open and
46 : : * close at stream start and close. We decided to use FileSet
47 : : * infrastructure as without that it deletes the files on the closure of the
48 : : * file and if we decide to keep stream files open across the start/stop stream
49 : : * then it will consume a lot of memory (more than 8K for each BufFile and
50 : : * there could be multiple such BufFiles as the subscriber could receive
51 : : * multiple start/stop streams for different transactions before getting the
52 : : * commit). Moreover, if we don't use FileSet then we also need to invent
53 : : * a new way to pass filenames to BufFile APIs so that we are allowed to open
54 : : * the file we desired across multiple stream-open calls for the same
55 : : * transaction.
56 : : *
57 : : * 2) Parallel apply workers.
58 : : *
59 : : * This approach is used when the user has set the subscription's streaming
60 : : * option as parallel. See logical/applyparallelworker.c for information about
61 : : * this approach.
62 : : *
63 : : * TWO_PHASE TRANSACTIONS
64 : : * ----------------------
65 : : * Two phase transactions are replayed at prepare and then committed or
66 : : * rolled back at commit prepared and rollback prepared respectively. It is
67 : : * possible to have a prepared transaction that arrives at the apply worker
68 : : * when the tablesync is busy doing the initial copy. In this case, the apply
69 : : * worker skips all the prepared operations [e.g. inserts] while the tablesync
70 : : * is still busy (see the condition of should_apply_changes_for_rel). The
71 : : * tablesync worker might not get such a prepared transaction because say it
72 : : * was prior to the initial consistent point but might have got some later
73 : : * commits. Now, the tablesync worker will exit without doing anything for the
74 : : * prepared transaction skipped by the apply worker as the sync location for it
75 : : * will be already ahead of the apply worker's current location. This would lead
76 : : * to an "empty prepare", because later when the apply worker does the commit
77 : : * prepare, there is nothing in it (the inserts were skipped earlier).
78 : : *
79 : : * To avoid this, and similar prepare confusions the subscription's two_phase
80 : : * commit is enabled only after the initial sync is over. The two_phase option
81 : : * has been implemented as a tri-state with values DISABLED, PENDING, and
82 : : * ENABLED.
83 : : *
84 : : * Even if the user specifies they want a subscription with two_phase = on,
85 : : * internally it will start with a tri-state of PENDING which only becomes
86 : : * ENABLED after all tablesync initializations are completed - i.e. when all
87 : : * tablesync workers have reached their READY state. In other words, the value
88 : : * PENDING is only a temporary state for subscription start-up.
89 : : *
90 : : * Until the two_phase is properly available (ENABLED) the subscription will
91 : : * behave as if two_phase = off. When the apply worker detects that all
92 : : * tablesyncs have become READY (while the tri-state was PENDING) it will
93 : : * restart the apply worker process. This happens in
94 : : * ProcessSyncingTablesForApply.
95 : : *
96 : : * When the (re-started) apply worker finds that all tablesyncs are READY for a
97 : : * two_phase tri-state of PENDING it start streaming messages with the
98 : : * two_phase option which in turn enables the decoding of two-phase commits at
99 : : * the publisher. Then, it updates the tri-state value from PENDING to ENABLED.
100 : : * Now, it is possible that during the time we have not enabled two_phase, the
101 : : * publisher (replication server) would have skipped some prepares but we
102 : : * ensure that such prepares are sent along with commit prepare, see
103 : : * ReorderBufferFinishPrepared.
104 : : *
105 : : * If the subscription has no tables then a two_phase tri-state PENDING is
106 : : * left unchanged. This lets the user still do an ALTER SUBSCRIPTION REFRESH
107 : : * PUBLICATION which might otherwise be disallowed (see below).
108 : : *
109 : : * If ever a user needs to be aware of the tri-state value, they can fetch it
110 : : * from the pg_subscription catalog (see column subtwophasestate).
111 : : *
112 : : * Finally, to avoid problems mentioned in previous paragraphs from any
113 : : * subsequent (not READY) tablesyncs (need to toggle two_phase option from 'on'
114 : : * to 'off' and then again back to 'on') there is a restriction for
115 : : * ALTER SUBSCRIPTION REFRESH PUBLICATION. This command is not permitted when
116 : : * the two_phase tri-state is ENABLED, except when copy_data = false.
117 : : *
118 : : * We can get prepare of the same GID more than once for the genuine cases
119 : : * where we have defined multiple subscriptions for publications on the same
120 : : * server and prepared transaction has operations on tables subscribed to those
121 : : * subscriptions. For such cases, if we use the GID sent by publisher one of
122 : : * the prepares will be successful and others will fail, in which case the
123 : : * server will send them again. Now, this can lead to a deadlock if user has
124 : : * set synchronous_standby_names for all the subscriptions on subscriber. To
125 : : * avoid such deadlocks, we generate a unique GID (consisting of the
126 : : * subscription oid and the xid of the prepared transaction) for each prepare
127 : : * transaction on the subscriber.
128 : : *
129 : : * FAILOVER
130 : : * ----------------------
131 : : * The logical slot on the primary can be synced to the standby by specifying
132 : : * failover = true when creating the subscription. Enabling failover allows us
133 : : * to smoothly transition to the promoted standby, ensuring that we can
134 : : * subscribe to the new primary without losing any data.
135 : : *
136 : : * RETAIN DEAD TUPLES
137 : : * ----------------------
138 : : * Each apply worker that enabled retain_dead_tuples option maintains a
139 : : * non-removable transaction ID (oldest_nonremovable_xid) in shared memory to
140 : : * prevent dead rows from being removed prematurely when the apply worker still
141 : : * needs them to detect update_deleted conflicts. Additionally, this helps to
142 : : * retain the required commit_ts module information, which further helps to
143 : : * detect update_origin_differs and delete_origin_differs conflicts reliably, as
144 : : * otherwise, vacuum freeze could remove the required information.
145 : : *
146 : : * The logical replication launcher manages an internal replication slot named
147 : : * "pg_conflict_detection". It asynchronously aggregates the non-removable
148 : : * transaction ID from all apply workers to determine the appropriate xmin for
149 : : * the slot, thereby retaining necessary tuples.
150 : : *
151 : : * The non-removable transaction ID in the apply worker is advanced to the
152 : : * oldest running transaction ID once all concurrent transactions on the
153 : : * publisher have been applied and flushed locally. The process involves:
154 : : *
155 : : * - RDT_GET_CANDIDATE_XID:
156 : : * Call GetOldestActiveTransactionId() to take oldestRunningXid as the
157 : : * candidate xid.
158 : : *
159 : : * - RDT_REQUEST_PUBLISHER_STATUS:
160 : : * Send a message to the walsender requesting the publisher status, which
161 : : * includes the latest WAL insert position and information about
162 : : * transactions that are in the commit phase.
163 : : *
164 : : * - RDT_WAIT_FOR_PUBLISHER_STATUS:
165 : : * Wait for the status from the walsender. After receiving the first status,
166 : : * do not proceed if there are concurrent remote transactions that are still
167 : : * in the commit phase. These transactions might have been assigned an
168 : : * earlier commit timestamp but have not yet written the commit WAL record.
169 : : * Continue to request the publisher status (RDT_REQUEST_PUBLISHER_STATUS)
170 : : * until all these transactions have completed.
171 : : *
172 : : * - RDT_WAIT_FOR_LOCAL_FLUSH:
173 : : * Advance the non-removable transaction ID if the current flush location has
174 : : * reached or surpassed the last received WAL position.
175 : : *
176 : : * - RDT_STOP_CONFLICT_INFO_RETENTION:
177 : : * This phase is required only when max_retention_duration is defined. We
178 : : * enter this phase if the wait time in either the
179 : : * RDT_WAIT_FOR_PUBLISHER_STATUS or RDT_WAIT_FOR_LOCAL_FLUSH phase exceeds
180 : : * configured max_retention_duration. In this phase,
181 : : * pg_subscription.subretentionactive is updated to false within a new
182 : : * transaction, and oldest_nonremovable_xid is set to InvalidTransactionId.
183 : : *
184 : : * - RDT_RESUME_CONFLICT_INFO_RETENTION:
185 : : * This phase is required only when max_retention_duration is defined. We
186 : : * enter this phase if the retention was previously stopped, and the time
187 : : * required to advance the non-removable transaction ID in the
188 : : * RDT_WAIT_FOR_LOCAL_FLUSH phase has decreased to within acceptable limits
189 : : * (or if max_retention_duration is set to 0). During this phase,
190 : : * pg_subscription.subretentionactive is updated to true within a new
191 : : * transaction, and the worker will be restarted.
192 : : *
193 : : * The overall state progression is: GET_CANDIDATE_XID ->
194 : : * REQUEST_PUBLISHER_STATUS -> WAIT_FOR_PUBLISHER_STATUS -> (loop to
195 : : * REQUEST_PUBLISHER_STATUS till concurrent remote transactions end) ->
196 : : * WAIT_FOR_LOCAL_FLUSH -> loop back to GET_CANDIDATE_XID.
197 : : *
198 : : * Retaining the dead tuples for this period is sufficient for ensuring
199 : : * eventual consistency using last-update-wins strategy, as dead tuples are
200 : : * useful for detecting conflicts only during the application of concurrent
201 : : * transactions from remote nodes. After applying and flushing all remote
202 : : * transactions that occurred concurrently with the tuple DELETE, any
203 : : * subsequent UPDATE from a remote node should have a later timestamp. In such
204 : : * cases, it is acceptable to detect an update_missing scenario and convert the
205 : : * UPDATE to an INSERT when applying it. But, for concurrent remote
206 : : * transactions with earlier timestamps than the DELETE, detecting
207 : : * update_deleted is necessary, as the UPDATEs in remote transactions should be
208 : : * ignored if their timestamp is earlier than that of the dead tuples.
209 : : *
210 : : * Note that advancing the non-removable transaction ID is not supported if the
211 : : * publisher is also a physical standby. This is because the logical walsender
212 : : * on the standby can only get the WAL replay position but there may be more
213 : : * WALs that are being replicated from the primary and those WALs could have
214 : : * earlier commit timestamp.
215 : : *
216 : : * Similarly, when the publisher has subscribed to another publisher,
217 : : * information necessary for conflict detection cannot be retained for
218 : : * changes from origins other than the publisher. This is because publisher
219 : : * lacks the information on concurrent transactions of other publishers to
220 : : * which it subscribes. As the information on concurrent transactions is
221 : : * unavailable beyond subscriber's immediate publishers, the non-removable
222 : : * transaction ID might be advanced prematurely before changes from other
223 : : * origins have been fully applied.
224 : : *
225 : : * XXX Retaining information for changes from other origins might be possible
226 : : * by requesting the subscription on that origin to enable retain_dead_tuples
227 : : * and fetching the conflict detection slot.xmin along with the publisher's
228 : : * status. In the RDT_WAIT_FOR_PUBLISHER_STATUS phase, the apply worker could
229 : : * wait for the remote slot's xmin to reach the oldest active transaction ID,
230 : : * ensuring that all transactions from other origins have been applied on the
231 : : * publisher, thereby getting the latest WAL position that includes all
232 : : * concurrent changes. However, this approach may impact performance, so it
233 : : * might not worth the effort.
234 : : *
235 : : * XXX It seems feasible to get the latest commit's WAL location from the
236 : : * publisher and wait till that is applied. However, we can't do that
237 : : * because commit timestamps can regress as a commit with a later LSN is not
238 : : * guaranteed to have a later timestamp than those with earlier LSNs. Having
239 : : * said that, even if that is possible, it won't improve performance much as
240 : : * the apply always lag and moves slowly as compared with the transactions
241 : : * on the publisher.
242 : : *-------------------------------------------------------------------------
243 : : */
244 : :
245 : : #include "postgres.h"
246 : :
247 : : #include <sys/stat.h>
248 : : #include <unistd.h>
249 : :
250 : : #include "access/genam.h"
251 : : #include "access/commit_ts.h"
252 : : #include "access/table.h"
253 : : #include "access/tableam.h"
254 : : #include "access/tupconvert.h"
255 : : #include "access/twophase.h"
256 : : #include "access/xact.h"
257 : : #include "catalog/indexing.h"
258 : : #include "catalog/pg_inherits.h"
259 : : #include "catalog/pg_subscription.h"
260 : : #include "catalog/pg_subscription_rel.h"
261 : : #include "commands/subscriptioncmds.h"
262 : : #include "commands/tablecmds.h"
263 : : #include "commands/trigger.h"
264 : : #include "executor/executor.h"
265 : : #include "executor/execPartition.h"
266 : : #include "libpq/pqformat.h"
267 : : #include "miscadmin.h"
268 : : #include "optimizer/optimizer.h"
269 : : #include "parser/parse_relation.h"
270 : : #include "pgstat.h"
271 : : #include "port/pg_bitutils.h"
272 : : #include "postmaster/bgworker.h"
273 : : #include "postmaster/interrupt.h"
274 : : #include "postmaster/walwriter.h"
275 : : #include "replication/conflict.h"
276 : : #include "replication/logicallauncher.h"
277 : : #include "replication/logicalproto.h"
278 : : #include "replication/logicalrelation.h"
279 : : #include "replication/logicalworker.h"
280 : : #include "replication/origin.h"
281 : : #include "replication/slot.h"
282 : : #include "replication/walreceiver.h"
283 : : #include "replication/worker_internal.h"
284 : : #include "rewrite/rewriteHandler.h"
285 : : #include "storage/buffile.h"
286 : : #include "storage/ipc.h"
287 : : #include "storage/latch.h"
288 : : #include "storage/lmgr.h"
289 : : #include "storage/procarray.h"
290 : : #include "tcop/tcopprot.h"
291 : : #include "utils/acl.h"
292 : : #include "utils/guc.h"
293 : : #include "utils/injection_point.h"
294 : : #include "utils/inval.h"
295 : : #include "utils/lsyscache.h"
296 : : #include "utils/memutils.h"
297 : : #include "utils/pg_lsn.h"
298 : : #include "utils/rel.h"
299 : : #include "utils/rls.h"
300 : : #include "utils/snapmgr.h"
301 : : #include "utils/syscache.h"
302 : : #include "utils/usercontext.h"
303 : : #include "utils/wait_event.h"
304 : :
305 : : #define NAPTIME_PER_CYCLE 1000 /* max sleep time between cycles (1s) */
306 : :
307 : : typedef struct FlushPosition
308 : : {
309 : : dlist_node node;
310 : : XLogRecPtr local_end;
311 : : XLogRecPtr remote_end;
312 : : } FlushPosition;
313 : :
314 : : static dlist_head lsn_mapping = DLIST_STATIC_INIT(lsn_mapping);
315 : :
316 : : typedef struct ApplyExecutionData
317 : : {
318 : : EState *estate; /* executor state, used to track resources */
319 : :
320 : : LogicalRepRelMapEntry *targetRel; /* replication target rel */
321 : : ResultRelInfo *targetRelInfo; /* ResultRelInfo for same */
322 : :
323 : : /* These fields are used when the target relation is partitioned: */
324 : : ModifyTableState *mtstate; /* dummy ModifyTable state */
325 : : PartitionTupleRouting *proute; /* partition routing info */
326 : : } ApplyExecutionData;
327 : :
328 : : /*
329 : : * Context describing the remote transaction whose changes are currently
330 : : * being applied, and the change within it.
331 : : *
332 : : * The remote transaction information (remote_xid and finish_lsn) is set when
333 : : * the transaction's changes begin to be applied. finish_lsn is invalid when
334 : : * the final LSN of the remote transaction is not yet known (e.g. while
335 : : * streaming an in-progress transaction).
336 : : *
337 : : * The remaining fields describe the individual change being applied and are
338 : : * used only for error context reporting.
339 : : */
340 : : typedef struct ApplyRemoteCtx
341 : : {
342 : : LogicalRepMsgType command; /* 0 if invalid */
343 : : LogicalRepRelMapEntry *rel;
344 : :
345 : : /* Remote node information */
346 : : int remote_attnum; /* -1 if invalid */
347 : : TransactionId remote_xid;
348 : : XLogRecPtr finish_lsn;
349 : : char *origin_name;
350 : : } ApplyRemoteCtx;
351 : :
352 : : /*
353 : : * The action to be taken for the changes in the transaction.
354 : : *
355 : : * TRANS_LEADER_APPLY:
356 : : * This action means that we are in the leader apply worker or table sync
357 : : * worker. The changes of the transaction are either directly applied or
358 : : * are read from temporary files (for streaming transactions) and then
359 : : * applied by the worker.
360 : : *
361 : : * TRANS_LEADER_SERIALIZE:
362 : : * This action means that we are in the leader apply worker or table sync
363 : : * worker. Changes are written to temporary files and then applied when the
364 : : * final commit arrives.
365 : : *
366 : : * TRANS_LEADER_SEND_TO_PARALLEL:
367 : : * This action means that we are in the leader apply worker and need to send
368 : : * the changes to the parallel apply worker.
369 : : *
370 : : * TRANS_LEADER_PARTIAL_SERIALIZE:
371 : : * This action means that we are in the leader apply worker and have sent some
372 : : * changes directly to the parallel apply worker and the remaining changes are
373 : : * serialized to a file, due to timeout while sending data. The parallel apply
374 : : * worker will apply these serialized changes when the final commit arrives.
375 : : *
376 : : * We can't use TRANS_LEADER_SERIALIZE for this case because, in addition to
377 : : * serializing changes, the leader worker also needs to serialize the
378 : : * STREAM_XXX message to a file, and wait for the parallel apply worker to
379 : : * finish the transaction when processing the transaction finish command. So
380 : : * this new action was introduced to keep the code and logic clear.
381 : : *
382 : : * TRANS_PARALLEL_APPLY:
383 : : * This action means that we are in the parallel apply worker and changes of
384 : : * the transaction are applied directly by the worker.
385 : : */
386 : : typedef enum
387 : : {
388 : : /* The action for non-streaming transactions. */
389 : : TRANS_LEADER_APPLY,
390 : :
391 : : /* Actions for streaming transactions. */
392 : : TRANS_LEADER_SERIALIZE,
393 : : TRANS_LEADER_SEND_TO_PARALLEL,
394 : : TRANS_LEADER_PARTIAL_SERIALIZE,
395 : : TRANS_PARALLEL_APPLY,
396 : : } TransApplyAction;
397 : :
398 : : /*
399 : : * The phases involved in advancing the non-removable transaction ID.
400 : : *
401 : : * See comments atop worker.c for details of the transition between these
402 : : * phases.
403 : : */
404 : : typedef enum
405 : : {
406 : : RDT_GET_CANDIDATE_XID,
407 : : RDT_REQUEST_PUBLISHER_STATUS,
408 : : RDT_WAIT_FOR_PUBLISHER_STATUS,
409 : : RDT_WAIT_FOR_LOCAL_FLUSH,
410 : : RDT_STOP_CONFLICT_INFO_RETENTION,
411 : : RDT_RESUME_CONFLICT_INFO_RETENTION,
412 : : } RetainDeadTuplesPhase;
413 : :
414 : : /*
415 : : * Critical information for managing phase transitions within the
416 : : * RetainDeadTuplesPhase.
417 : : */
418 : : typedef struct RetainDeadTuplesData
419 : : {
420 : : RetainDeadTuplesPhase phase; /* current phase */
421 : : XLogRecPtr remote_lsn; /* WAL insert position on the publisher */
422 : :
423 : : /*
424 : : * Oldest transaction ID that was in the commit phase on the publisher.
425 : : * Use FullTransactionId to prevent issues with transaction ID wraparound,
426 : : * where a new remote_oldestxid could falsely appear to originate from the
427 : : * past and block advancement.
428 : : */
429 : : FullTransactionId remote_oldestxid;
430 : :
431 : : /*
432 : : * Next transaction ID to be assigned on the publisher. Use
433 : : * FullTransactionId for consistency and to allow straightforward
434 : : * comparisons with remote_oldestxid.
435 : : */
436 : : FullTransactionId remote_nextxid;
437 : :
438 : : TimestampTz reply_time; /* when the publisher responds with status */
439 : :
440 : : /*
441 : : * Publisher transaction ID that must be awaited to complete before
442 : : * entering the final phase (RDT_WAIT_FOR_LOCAL_FLUSH). Use
443 : : * FullTransactionId for the same reason as remote_nextxid.
444 : : */
445 : : FullTransactionId remote_wait_for;
446 : :
447 : : TransactionId candidate_xid; /* candidate for the non-removable
448 : : * transaction ID */
449 : : TimestampTz flushpos_update_time; /* when the remote flush position was
450 : : * updated in final phase
451 : : * (RDT_WAIT_FOR_LOCAL_FLUSH) */
452 : :
453 : : long table_sync_wait_time; /* time spent waiting for table sync
454 : : * to finish */
455 : :
456 : : /*
457 : : * The following fields are used to determine the timing for the next
458 : : * round of transaction ID advancement.
459 : : */
460 : : TimestampTz last_recv_time; /* when the last message was received */
461 : : TimestampTz candidate_xid_time; /* when the candidate_xid is decided */
462 : : int xid_advance_interval; /* how much time (ms) to wait before
463 : : * attempting to advance the
464 : : * non-removable transaction ID */
465 : : } RetainDeadTuplesData;
466 : :
467 : : /*
468 : : * The minimum (100ms) and maximum (3 minutes) intervals for advancing
469 : : * non-removable transaction IDs. The maximum interval is a bit arbitrary but
470 : : * is sufficient to not cause any undue network traffic.
471 : : */
472 : : #define MIN_XID_ADVANCE_INTERVAL 100
473 : : #define MAX_XID_ADVANCE_INTERVAL 180000
474 : :
475 : : /* Context of the remote transaction being applied */
476 : : static ApplyRemoteCtx remote_ctx =
477 : : {
478 : : .command = 0,
479 : : .rel = NULL,
480 : : .remote_attnum = -1,
481 : : .remote_xid = InvalidTransactionId,
482 : : .finish_lsn = InvalidXLogRecPtr,
483 : : .origin_name = NULL,
484 : : };
485 : :
486 : : ErrorContextCallback *apply_error_context_stack = NULL;
487 : :
488 : : MemoryContext ApplyMessageContext = NULL;
489 : : MemoryContext ApplyContext = NULL;
490 : :
491 : : /* per stream context for streaming transactions */
492 : : static MemoryContext LogicalStreamingContext = NULL;
493 : :
494 : : WalReceiverConn *LogRepWorkerWalRcvConn = NULL;
495 : :
496 : : Subscription *MySubscription = NULL;
497 : : char *MySubscriptionConninfo = NULL;
498 : : static bool MySubscriptionValid = false;
499 : :
500 : : static List *on_commit_wakeup_workers_subids = NIL;
501 : :
502 : : bool in_remote_transaction = false;
503 : :
504 : : /* fields valid only when processing streamed transaction */
505 : : static bool in_streamed_transaction = false;
506 : :
507 : : static TransactionId stream_xid = InvalidTransactionId;
508 : :
509 : : /*
510 : : * The number of changes applied by parallel apply worker during one streaming
511 : : * block.
512 : : */
513 : : static uint32 parallel_stream_nchanges = 0;
514 : :
515 : : /* Are we initializing an apply worker? */
516 : : bool InitializingApplyWorker = false;
517 : :
518 : : /*
519 : : * We enable skipping all data modification changes (INSERT, UPDATE, etc.) for
520 : : * the subscription if the remote transaction's finish LSN matches the subskiplsn.
521 : : * Once we start skipping changes, we don't stop it until we skip all changes of
522 : : * the transaction even if pg_subscription is updated and MySubscription->skiplsn
523 : : * gets changed or reset during that. Also, in streaming transaction cases (streaming = on),
524 : : * we don't skip receiving and spooling the changes since we decide whether or not
525 : : * to skip applying the changes when starting to apply changes. The subskiplsn is
526 : : * cleared after successfully skipping the transaction or applying non-empty
527 : : * transaction. The latter prevents the mistakenly specified subskiplsn from
528 : : * being left. Note that we cannot skip the streaming transactions when using
529 : : * parallel apply workers because we cannot get the finish LSN before applying
530 : : * the changes. So, we don't start parallel apply worker when finish LSN is set
531 : : * by the user.
532 : : */
533 : : static XLogRecPtr skip_xact_finish_lsn = InvalidXLogRecPtr;
534 : : #define is_skipping_changes() (unlikely(XLogRecPtrIsValid(skip_xact_finish_lsn)))
535 : :
536 : : /* BufFile handle of the current streaming file */
537 : : static BufFile *stream_fd = NULL;
538 : :
539 : : /*
540 : : * The remote WAL position that has been applied and flushed locally. We record
541 : : * and use this information both while sending feedback to the server and
542 : : * advancing oldest_nonremovable_xid.
543 : : */
544 : : static XLogRecPtr last_flushpos = InvalidXLogRecPtr;
545 : :
546 : : typedef struct SubXactInfo
547 : : {
548 : : TransactionId xid; /* XID of the subxact */
549 : : int fileno; /* file number in the buffile */
550 : : pgoff_t offset; /* offset in the file */
551 : : } SubXactInfo;
552 : :
553 : : /* Sub-transaction data for the current streaming transaction */
554 : : typedef struct ApplySubXactData
555 : : {
556 : : uint32 nsubxacts; /* number of sub-transactions */
557 : : uint32 nsubxacts_max; /* current capacity of subxacts */
558 : : TransactionId subxact_last; /* xid of the last sub-transaction */
559 : : SubXactInfo *subxacts; /* sub-xact offset in changes file */
560 : : } ApplySubXactData;
561 : :
562 : : static ApplySubXactData subxact_data = {0, 0, InvalidTransactionId, NULL};
563 : :
564 : : static inline void subxact_filename(char *path, Oid subid, TransactionId xid);
565 : : static inline void changes_filename(char *path, Oid subid, TransactionId xid);
566 : :
567 : : /*
568 : : * Information about subtransactions of a given toplevel transaction.
569 : : */
570 : : static void subxact_info_write(Oid subid, TransactionId xid);
571 : : static void subxact_info_read(Oid subid, TransactionId xid);
572 : : static void subxact_info_add(TransactionId xid);
573 : : static inline void cleanup_subxact_info(void);
574 : :
575 : : /*
576 : : * Serialize and deserialize changes for a toplevel transaction.
577 : : */
578 : : static void stream_open_file(Oid subid, TransactionId xid,
579 : : bool first_segment);
580 : : static void stream_write_change(char action, StringInfo s);
581 : : static void stream_open_and_write_change(TransactionId xid, char action, StringInfo s);
582 : : static void stream_close_file(void);
583 : :
584 : : static void send_feedback(XLogRecPtr recvpos, bool force, bool requestReply);
585 : :
586 : : static void maybe_advance_nonremovable_xid(RetainDeadTuplesData *rdt_data,
587 : : bool status_received);
588 : : static bool can_advance_nonremovable_xid(RetainDeadTuplesData *rdt_data);
589 : : static void process_rdt_phase_transition(RetainDeadTuplesData *rdt_data,
590 : : bool status_received);
591 : : static void get_candidate_xid(RetainDeadTuplesData *rdt_data);
592 : : static void request_publisher_status(RetainDeadTuplesData *rdt_data);
593 : : static void wait_for_publisher_status(RetainDeadTuplesData *rdt_data,
594 : : bool status_received);
595 : : static void wait_for_local_flush(RetainDeadTuplesData *rdt_data);
596 : : static bool should_stop_conflict_info_retention(RetainDeadTuplesData *rdt_data);
597 : : static void stop_conflict_info_retention(RetainDeadTuplesData *rdt_data);
598 : : static void resume_conflict_info_retention(RetainDeadTuplesData *rdt_data);
599 : : static bool update_retention_status(bool active);
600 : : static void reset_retention_data_fields(RetainDeadTuplesData *rdt_data);
601 : : static void adjust_xid_advance_interval(RetainDeadTuplesData *rdt_data,
602 : : bool new_xid_found);
603 : :
604 : : static void apply_worker_exit(void);
605 : :
606 : : static void apply_handle_commit_internal(LogicalRepCommitData *commit_data);
607 : : static void apply_handle_insert_internal(ApplyExecutionData *edata,
608 : : ResultRelInfo *relinfo,
609 : : TupleTableSlot *remoteslot);
610 : : static void apply_handle_update_internal(ApplyExecutionData *edata,
611 : : ResultRelInfo *relinfo,
612 : : TupleTableSlot *remoteslot,
613 : : LogicalRepTupleData *newtup);
614 : : static void apply_handle_delete_internal(ApplyExecutionData *edata,
615 : : ResultRelInfo *relinfo,
616 : : TupleTableSlot *remoteslot,
617 : : LogicalRepRelMapEntry *relmapentry);
618 : : static bool FindReplTupleInLocalRel(ApplyExecutionData *edata, Relation localrel,
619 : : LogicalRepRelMapEntry *relmapentry,
620 : : TupleTableSlot *remoteslot,
621 : : TupleTableSlot **localslot);
622 : : static bool FindDeletedTupleInLocalRel(Relation localrel,
623 : : LogicalRepRelMapEntry *relmapentry,
624 : : TupleTableSlot *remoteslot,
625 : : TransactionId *delete_xid,
626 : : ReplOriginId *delete_origin,
627 : : TimestampTz *delete_time);
628 : : static void apply_handle_tuple_routing(ApplyExecutionData *edata,
629 : : TupleTableSlot *remoteslot,
630 : : LogicalRepTupleData *newtup,
631 : : CmdType operation);
632 : :
633 : : /* Functions for skipping changes */
634 : : static void maybe_start_skipping_changes(XLogRecPtr finish_lsn);
635 : : static void stop_skipping_changes(void);
636 : : static void clear_subscription_skip_lsn(XLogRecPtr finish_lsn);
637 : :
638 : : /* Functions to maintain the context of the remote transaction being applied */
639 : : static inline void set_remote_transaction_info(TransactionId xid, XLogRecPtr lsn);
640 : : static inline void reset_apply_remote_context(void);
641 : :
642 : : static TransApplyAction get_transaction_apply_action(TransactionId xid,
643 : : ParallelApplyWorkerInfo **winfo);
644 : :
645 : : static void set_wal_receiver_timeout(void);
646 : :
647 : : static void on_exit_clear_xact_state(int code, Datum arg);
648 : :
649 : : /*
650 : : * Form the origin name for the subscription.
651 : : *
652 : : * This is a common function for tablesync and other workers. Tablesync workers
653 : : * must pass a valid relid. Other callers must pass relid = InvalidOid.
654 : : *
655 : : * Return the name in the supplied buffer.
656 : : */
657 : : void
658 : 1606 : ReplicationOriginNameForLogicalRep(Oid suboid, Oid relid,
659 : : char *originname, Size szoriginname)
660 : : {
661 [ + + ]: 1606 : if (OidIsValid(relid))
662 : : {
663 : : /* Replication origin name for tablesync workers. */
664 : 840 : snprintf(originname, szoriginname, "pg_%u_%u", suboid, relid);
665 : : }
666 : : else
667 : : {
668 : : /* Replication origin name for non-tablesync workers. */
669 : 766 : snprintf(originname, szoriginname, "pg_%u", suboid);
670 : : }
671 : 1606 : }
672 : :
673 : : /*
674 : : * Should this worker apply changes for given relation.
675 : : *
676 : : * This is mainly needed for initial relation data sync as that runs in
677 : : * separate worker process running in parallel and we need some way to skip
678 : : * changes coming to the leader apply worker during the sync of a table.
679 : : *
680 : : * Note we need to do smaller or equals comparison for SYNCDONE state because
681 : : * it might hold position of end of initial slot consistent point WAL
682 : : * record + 1 (ie start of next record) and next record can be COMMIT of
683 : : * transaction we are now processing (which is what we set the finish LSN of
684 : : * the remote transaction context to in apply_handle_begin).
685 : : *
686 : : * Note that for streaming transactions that are being applied in the parallel
687 : : * apply worker, we disallow applying changes if the target table in the
688 : : * subscription is not in the READY state, because we cannot decide whether to
689 : : * apply the change as we won't know the finish LSN of the transaction by
690 : : * that time.
691 : : *
692 : : * We already checked this in pa_can_start() before assigning the
693 : : * streaming transaction to the parallel worker, but it also needs to be
694 : : * checked here because if the user executes ALTER SUBSCRIPTION ... REFRESH
695 : : * PUBLICATION in parallel, the new table can be added to pg_subscription_rel
696 : : * while applying this transaction.
697 : : */
698 : : static bool
699 : 169050 : should_apply_changes_for_rel(LogicalRepRelMapEntry *rel)
700 : : {
701 [ - + + - : 169050 : switch (MyLogicalRepWorker->type)
- - ]
702 : : {
703 : 0 : case WORKERTYPE_TABLESYNC:
704 : 0 : return MyLogicalRepWorker->relid == rel->localreloid;
705 : :
706 : 68580 : case WORKERTYPE_PARALLEL_APPLY:
707 : : /* We don't synchronize rel's that are in unknown state. */
708 [ - + ]: 68580 : if (rel->state != SUBREL_STATE_READY &&
709 [ # # ]: 0 : rel->state != SUBREL_STATE_UNKNOWN)
710 [ # # ]: 0 : ereport(ERROR,
711 : : (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
712 : : errmsg("logical replication parallel apply worker for subscription \"%s\" will stop",
713 : : MySubscription->name),
714 : : errdetail("Cannot handle streamed replication transactions using parallel apply workers until all tables have been synchronized.")));
715 : :
716 : 68580 : return rel->state == SUBREL_STATE_READY;
717 : :
718 : 100470 : case WORKERTYPE_APPLY:
719 [ + + ]: 100533 : return (rel->state == SUBREL_STATE_READY ||
720 [ + + ]: 63 : (rel->state == SUBREL_STATE_SYNCDONE &&
721 [ + - ]: 15 : rel->statelsn <= remote_ctx.finish_lsn));
722 : :
723 : 0 : case WORKERTYPE_SEQUENCESYNC:
724 : : /* Should never happen. */
725 [ # # ]: 0 : elog(ERROR, "sequence synchronization worker is not expected to apply changes");
726 : : break;
727 : :
728 : 0 : case WORKERTYPE_UNKNOWN:
729 : : /* Should never happen. */
730 [ # # ]: 0 : elog(ERROR, "Unknown worker type");
731 : : }
732 : :
733 : 0 : return false; /* dummy for compiler */
734 : : }
735 : :
736 : : /*
737 : : * Begin one step (one INSERT, UPDATE, etc) of a replication transaction.
738 : : *
739 : : * Start a transaction, if this is the first step (else we keep using the
740 : : * existing transaction).
741 : : * Also provide a global snapshot and ensure we run in ApplyMessageContext.
742 : : */
743 : : static void
744 : 169514 : begin_replication_step(void)
745 : : {
746 : 169514 : SetCurrentStatementStartTimestamp();
747 : :
748 [ + + ]: 169514 : if (!IsTransactionState())
749 : : {
750 : 1027 : StartTransactionCommand();
751 : 1027 : maybe_reread_subscription();
752 : : }
753 : :
754 : 169511 : PushActiveSnapshot(GetTransactionSnapshot());
755 : :
756 : 169511 : MemoryContextSwitchTo(ApplyMessageContext);
757 : 169511 : }
758 : :
759 : : /*
760 : : * Finish up one step of a replication transaction.
761 : : * Callers of begin_replication_step() must also call this.
762 : : *
763 : : * We don't close out the transaction here, but we should increment
764 : : * the command counter to make the effects of this step visible.
765 : : */
766 : : static void
767 : 169423 : end_replication_step(void)
768 : : {
769 : 169423 : PopActiveSnapshot();
770 : :
771 : 169423 : CommandCounterIncrement();
772 : 169423 : }
773 : :
774 : : /*
775 : : * Handle streamed transactions for both the leader apply worker and the
776 : : * parallel apply workers.
777 : : *
778 : : * In the streaming case (receiving a block of the streamed transaction), for
779 : : * serialize mode, simply redirect it to a file for the proper toplevel
780 : : * transaction, and for parallel mode, the leader apply worker will send the
781 : : * changes to parallel apply workers and the parallel apply worker will define
782 : : * savepoints if needed. (LOGICAL_REP_MSG_RELATION or LOGICAL_REP_MSG_TYPE
783 : : * messages will be applied by both leader apply worker and parallel apply
784 : : * workers).
785 : : *
786 : : * Returns true for streamed transactions (when the change is either serialized
787 : : * to file or sent to parallel apply worker), false otherwise (regular mode or
788 : : * needs to be processed by parallel apply worker).
789 : : *
790 : : * Exception: If the message being processed is LOGICAL_REP_MSG_RELATION
791 : : * or LOGICAL_REP_MSG_TYPE, return false even if the message needs to be sent
792 : : * to a parallel apply worker.
793 : : */
794 : : static bool
795 : 345454 : handle_streamed_transaction(LogicalRepMsgType action, StringInfo s)
796 : : {
797 : : TransactionId current_xid;
798 : : ParallelApplyWorkerInfo *winfo;
799 : : TransApplyAction apply_action;
800 : : StringInfoData original_msg;
801 : :
802 : 345454 : apply_action = get_transaction_apply_action(stream_xid, &winfo);
803 : :
804 : : /* not in streaming mode */
805 [ + + ]: 345454 : if (apply_action == TRANS_LEADER_APPLY)
806 : 100935 : return false;
807 : :
808 : : Assert(TransactionIdIsValid(stream_xid));
809 : :
810 : : /*
811 : : * The parallel apply worker needs the xid in this message to decide
812 : : * whether to define a savepoint, so save the original message that has
813 : : * not moved the cursor after the xid. We will serialize this message to a
814 : : * file in PARTIAL_SERIALIZE mode.
815 : : */
816 : 244519 : original_msg = *s;
817 : :
818 : : /*
819 : : * We should have received XID of the subxact as the first part of the
820 : : * message, so extract it.
821 : : */
822 : 244519 : current_xid = pq_getmsgint(s, 4);
823 : :
824 [ - + ]: 244519 : if (!TransactionIdIsValid(current_xid))
825 [ # # ]: 0 : ereport(ERROR,
826 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
827 : : errmsg_internal("invalid transaction ID in streamed replication transaction")));
828 : :
829 [ + + + + : 244519 : switch (apply_action)
- ]
830 : : {
831 : 102513 : case TRANS_LEADER_SERIALIZE:
832 : : Assert(stream_fd);
833 : :
834 : : /* Add the new subxact to the array (unless already there). */
835 : 102513 : subxact_info_add(current_xid);
836 : :
837 : : /* Write the change to the current file */
838 : 102513 : stream_write_change(action, s);
839 : 102513 : return true;
840 : :
841 : 68390 : case TRANS_LEADER_SEND_TO_PARALLEL:
842 : : Assert(winfo);
843 : :
844 : : /*
845 : : * XXX The publisher side doesn't always send relation/type update
846 : : * messages after the streaming transaction, so also update the
847 : : * relation/type in leader apply worker. See function
848 : : * cleanup_rel_sync_cache.
849 : : */
850 [ + - ]: 68390 : if (pa_send_data(winfo, s->len, s->data))
851 [ + + + - ]: 68390 : return (action != LOGICAL_REP_MSG_RELATION &&
852 : : action != LOGICAL_REP_MSG_TYPE);
853 : :
854 : : /*
855 : : * Switch to serialize mode when we are not able to send the
856 : : * change to parallel apply worker.
857 : : */
858 : 0 : pa_switch_to_partial_serialize(winfo, false);
859 : :
860 : : pg_fallthrough;
861 : 5006 : case TRANS_LEADER_PARTIAL_SERIALIZE:
862 : 5006 : stream_write_change(action, &original_msg);
863 : :
864 : : /* Same reason as TRANS_LEADER_SEND_TO_PARALLEL case. */
865 [ + + + - ]: 5006 : return (action != LOGICAL_REP_MSG_RELATION &&
866 : : action != LOGICAL_REP_MSG_TYPE);
867 : :
868 : 68610 : case TRANS_PARALLEL_APPLY:
869 : 68610 : parallel_stream_nchanges += 1;
870 : :
871 : : /* Define a savepoint for a subxact if needed. */
872 : 68610 : pa_start_subtrans(current_xid, stream_xid);
873 : 68610 : return false;
874 : :
875 : 0 : default:
876 [ # # ]: 0 : elog(ERROR, "unexpected apply action: %d", (int) apply_action);
877 : : return false; /* silence compiler warning */
878 : : }
879 : : }
880 : :
881 : : /*
882 : : * Executor state preparation for evaluation of constraint expressions,
883 : : * indexes and triggers for the specified relation.
884 : : *
885 : : * Note that the caller must open and close any indexes to be updated.
886 : : */
887 : : static ApplyExecutionData *
888 : 168970 : create_edata_for_relation(LogicalRepRelMapEntry *rel)
889 : : {
890 : : ApplyExecutionData *edata;
891 : : EState *estate;
892 : : RangeTblEntry *rte;
893 : 168970 : List *perminfos = NIL;
894 : : ResultRelInfo *resultRelInfo;
895 : :
896 : 168970 : edata = palloc0_object(ApplyExecutionData);
897 : 168970 : edata->targetRel = rel;
898 : :
899 : 168970 : edata->estate = estate = CreateExecutorState();
900 : :
901 : 168970 : rte = makeNode(RangeTblEntry);
902 : 168970 : rte->rtekind = RTE_RELATION;
903 : 168970 : rte->relid = RelationGetRelid(rel->localrel);
904 : 168970 : rte->relkind = rel->localrel->rd_rel->relkind;
905 : 168970 : rte->rellockmode = AccessShareLock;
906 : :
907 : 168970 : addRTEPermissionInfo(&perminfos, rte);
908 : :
909 : 168970 : ExecInitRangeTable(estate, list_make1(rte), perminfos,
910 : : bms_make_singleton(1));
911 : :
912 : 168970 : edata->targetRelInfo = resultRelInfo = makeNode(ResultRelInfo);
913 : :
914 : : /*
915 : : * Use Relation opened by logicalrep_rel_open() instead of opening it
916 : : * again.
917 : : */
918 : 168970 : InitResultRelInfo(resultRelInfo, rel->localrel, 1, NULL, 0);
919 : :
920 : : /*
921 : : * We put the ResultRelInfo in the es_opened_result_relations list, even
922 : : * though we don't populate the es_result_relations array. That's a bit
923 : : * bogus, but it's enough to make ExecGetTriggerResultRel() find them.
924 : : *
925 : : * ExecOpenIndices() is not called here either, each execution path doing
926 : : * an apply operation being responsible for that.
927 : : */
928 : 168970 : estate->es_opened_result_relations =
929 : 168970 : lappend(estate->es_opened_result_relations, resultRelInfo);
930 : :
931 : 168970 : estate->es_output_cid = GetCurrentCommandId(true);
932 : :
933 : : /* Prepare to catch AFTER triggers. */
934 : 168970 : AfterTriggerBeginQuery();
935 : :
936 : : /* other fields of edata remain NULL for now */
937 : :
938 : 168970 : return edata;
939 : : }
940 : :
941 : : /*
942 : : * Finish any operations related to the executor state created by
943 : : * create_edata_for_relation().
944 : : */
945 : : static void
946 : 168897 : finish_edata(ApplyExecutionData *edata)
947 : : {
948 : 168897 : EState *estate = edata->estate;
949 : :
950 : : /* Handle any queued AFTER triggers. */
951 : 168897 : AfterTriggerEndQuery(estate);
952 : :
953 : : /* Shut down tuple routing, if any was done. */
954 [ + + ]: 168897 : if (edata->proute)
955 : 74 : ExecCleanupTupleRouting(edata->mtstate, edata->proute);
956 : :
957 : : /*
958 : : * Close relations opened specifically for trigger targets. It might seem
959 : : * that we should call ExecCloseResultRelations() here, but we
960 : : * intentionally don't as that would close the rel we added to
961 : : * es_opened_result_relations above, which is wrong because we took no
962 : : * corresponding refcount. ExecCleanupTupleRouting() closes relations
963 : : * opened for tuple routing, while ExecCloseTrigTargetRelations() closes
964 : : * any relations we opened for AFTER triggers.
965 : : */
966 : 168897 : ExecCloseTrigTargetRelations(estate);
967 : :
968 : 168897 : ExecResetTupleTable(estate->es_tupleTable, false);
969 : 168897 : FreeExecutorState(estate);
970 : 168897 : pfree(edata);
971 : 168897 : }
972 : :
973 : : /*
974 : : * Executes default values for columns for which we can't map to remote
975 : : * relation columns.
976 : : *
977 : : * This allows us to support tables which have more columns on the downstream
978 : : * than on the upstream.
979 : : */
980 : : static void
981 : 96697 : slot_fill_defaults(LogicalRepRelMapEntry *rel, EState *estate,
982 : : TupleTableSlot *slot)
983 : : {
984 : 96697 : TupleDesc desc = RelationGetDescr(rel->localrel);
985 : 96697 : int num_phys_attrs = desc->natts;
986 : : int i;
987 : : int attnum,
988 : 96697 : num_defaults = 0;
989 : : int *defmap;
990 : : ExprState **defexprs;
991 : : ExprContext *econtext;
992 : :
993 [ + - ]: 96697 : econtext = GetPerTupleExprContext(estate);
994 : :
995 : : /* We got all the data via replication, no need to evaluate anything. */
996 [ + + ]: 96697 : if (num_phys_attrs == rel->remoterel.natts)
997 : 56552 : return;
998 : :
999 : 40145 : defmap = palloc_array(int, num_phys_attrs);
1000 : 40145 : defexprs = palloc_array(ExprState *, num_phys_attrs);
1001 : :
1002 : : Assert(rel->attrmap->maplen == num_phys_attrs);
1003 [ + + ]: 210663 : for (attnum = 0; attnum < num_phys_attrs; attnum++)
1004 : : {
1005 : 170518 : CompactAttribute *cattr = TupleDescCompactAttr(desc, attnum);
1006 : : Expr *defexpr;
1007 : :
1008 [ + - + + ]: 170518 : if (cattr->attisdropped || cattr->attgenerated)
1009 : 9 : continue;
1010 : :
1011 [ + + ]: 170509 : if (rel->attrmap->attnums[attnum] >= 0)
1012 : 92268 : continue;
1013 : :
1014 : 78241 : defexpr = (Expr *) build_column_default(rel->localrel, attnum + 1);
1015 : :
1016 [ + + ]: 78241 : if (defexpr != NULL)
1017 : : {
1018 : : /* Run the expression through planner */
1019 : 70131 : defexpr = expression_planner(defexpr);
1020 : :
1021 : : /* Initialize executable expression in copycontext */
1022 : 70131 : defexprs[num_defaults] = ExecInitExpr(defexpr, NULL);
1023 : 70131 : defmap[num_defaults] = attnum;
1024 : 70131 : num_defaults++;
1025 : : }
1026 : : }
1027 : :
1028 [ + + ]: 110276 : for (i = 0; i < num_defaults; i++)
1029 : 70131 : slot->tts_values[defmap[i]] =
1030 : 70131 : ExecEvalExpr(defexprs[i], econtext, &slot->tts_isnull[defmap[i]]);
1031 : : }
1032 : :
1033 : : /*
1034 : : * Store tuple data into slot.
1035 : : *
1036 : : * Incoming data can be either text or binary format.
1037 : : */
1038 : : static void
1039 : 168991 : slot_store_data(TupleTableSlot *slot, LogicalRepRelMapEntry *rel,
1040 : : LogicalRepTupleData *tupleData)
1041 : : {
1042 : 168991 : int natts = slot->tts_tupleDescriptor->natts;
1043 : : int i;
1044 : :
1045 : 168991 : ExecClearTuple(slot);
1046 : :
1047 : : /* Call the "in" function for each non-dropped, non-null attribute */
1048 : : Assert(natts == rel->attrmap->maplen);
1049 [ + + ]: 699676 : for (i = 0; i < natts; i++)
1050 : : {
1051 : 530685 : Form_pg_attribute att = TupleDescAttr(slot->tts_tupleDescriptor, i);
1052 : 530685 : int remoteattnum = rel->attrmap->attnums[i];
1053 : :
1054 [ + + + + ]: 530685 : if (!att->attisdropped && remoteattnum >= 0)
1055 : 323768 : {
1056 : : StringInfo colvalue;
1057 : :
1058 [ - + ]: 323768 : if (remoteattnum >= tupleData->ncols)
1059 [ # # ]: 0 : ereport(ERROR,
1060 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1061 : : errmsg_plural("logical replication column %d not found in tuple: only %d column received",
1062 : : "logical replication column %d not found in tuple: only %d columns received",
1063 : : tupleData->ncols,
1064 : : remoteattnum + 1, tupleData->ncols)));
1065 : :
1066 : 323768 : colvalue = &tupleData->colvalues[remoteattnum];
1067 : :
1068 : : /* Set attnum for error callback */
1069 : 323768 : remote_ctx.remote_attnum = remoteattnum;
1070 : :
1071 [ + + ]: 323768 : if (tupleData->colstatus[remoteattnum] == LOGICALREP_COLUMN_TEXT)
1072 : : {
1073 : : Oid typinput;
1074 : : Oid typioparam;
1075 : :
1076 : 163233 : getTypeInputInfo(att->atttypid, &typinput, &typioparam);
1077 : 326466 : slot->tts_values[i] =
1078 : 163233 : OidInputFunctionCall(typinput, colvalue->data,
1079 : : typioparam, att->atttypmod);
1080 : 163233 : slot->tts_isnull[i] = false;
1081 : : }
1082 [ + + ]: 160535 : else if (tupleData->colstatus[remoteattnum] == LOGICALREP_COLUMN_BINARY)
1083 : : {
1084 : : Oid typreceive;
1085 : : Oid typioparam;
1086 : :
1087 : : /*
1088 : : * In some code paths we may be asked to re-parse the same
1089 : : * tuple data. Reset the StringInfo's cursor so that works.
1090 : : */
1091 : 110193 : colvalue->cursor = 0;
1092 : :
1093 : 110193 : getTypeBinaryInputInfo(att->atttypid, &typreceive, &typioparam);
1094 : 220386 : slot->tts_values[i] =
1095 : 110193 : OidReceiveFunctionCall(typreceive, colvalue,
1096 : : typioparam, att->atttypmod);
1097 : :
1098 : : /* Trouble if it didn't eat the whole buffer */
1099 [ - + ]: 110193 : if (colvalue->cursor != colvalue->len)
1100 [ # # ]: 0 : ereport(ERROR,
1101 : : (errcode(ERRCODE_INVALID_BINARY_REPRESENTATION),
1102 : : errmsg("incorrect binary data format in logical replication column %d",
1103 : : remoteattnum + 1)));
1104 : 110193 : slot->tts_isnull[i] = false;
1105 : : }
1106 : : else
1107 : : {
1108 : : /*
1109 : : * NULL value from remote. (We don't expect to see
1110 : : * LOGICALREP_COLUMN_UNCHANGED here, but if we do, treat it as
1111 : : * NULL.)
1112 : : */
1113 : 50342 : slot->tts_values[i] = (Datum) 0;
1114 : 50342 : slot->tts_isnull[i] = true;
1115 : : }
1116 : :
1117 : : /* Reset attnum for error callback */
1118 : 323768 : remote_ctx.remote_attnum = -1;
1119 : : }
1120 : : else
1121 : : {
1122 : : /*
1123 : : * We assign NULL to dropped attributes and missing values
1124 : : * (missing values should be later filled using
1125 : : * slot_fill_defaults).
1126 : : */
1127 : 206917 : slot->tts_values[i] = (Datum) 0;
1128 : 206917 : slot->tts_isnull[i] = true;
1129 : : }
1130 : : }
1131 : :
1132 : 168991 : ExecStoreVirtualTuple(slot);
1133 : 168991 : }
1134 : :
1135 : : /*
1136 : : * Replace updated columns with data from the LogicalRepTupleData struct.
1137 : : * This is somewhat similar to heap_modify_tuple but also calls the type
1138 : : * input functions on the user data.
1139 : : *
1140 : : * "slot" is filled with a copy of the tuple in "srcslot", replacing
1141 : : * columns provided in "tupleData" and leaving others as-is.
1142 : : *
1143 : : * Caution: unreplaced pass-by-ref columns in "slot" will point into the
1144 : : * storage for "srcslot". This is OK for current usage, but someday we may
1145 : : * need to materialize "slot" at the end to make it independent of "srcslot".
1146 : : */
1147 : : static void
1148 : 31927 : slot_modify_data(TupleTableSlot *slot, TupleTableSlot *srcslot,
1149 : : LogicalRepRelMapEntry *rel,
1150 : : LogicalRepTupleData *tupleData)
1151 : : {
1152 : 31927 : int natts = slot->tts_tupleDescriptor->natts;
1153 : : int i;
1154 : :
1155 : : /* We'll fill "slot" with a virtual tuple, so we must start with ... */
1156 : 31927 : ExecClearTuple(slot);
1157 : :
1158 : : /*
1159 : : * Copy all the column data from srcslot, so that we'll have valid values
1160 : : * for unreplaced columns.
1161 : : */
1162 : : Assert(natts == srcslot->tts_tupleDescriptor->natts);
1163 : 31927 : slot_getallattrs(srcslot);
1164 : 31927 : memcpy(slot->tts_values, srcslot->tts_values, natts * sizeof(Datum));
1165 : 31927 : memcpy(slot->tts_isnull, srcslot->tts_isnull, natts * sizeof(bool));
1166 : :
1167 : : /* Call the "in" function for each replaced attribute */
1168 : : Assert(natts == rel->attrmap->maplen);
1169 [ + + ]: 159289 : for (i = 0; i < natts; i++)
1170 : : {
1171 : 127362 : Form_pg_attribute att = TupleDescAttr(slot->tts_tupleDescriptor, i);
1172 : 127362 : int remoteattnum = rel->attrmap->attnums[i];
1173 : :
1174 [ + + ]: 127362 : if (remoteattnum < 0)
1175 : 58519 : continue;
1176 : :
1177 [ - + ]: 68843 : if (remoteattnum >= tupleData->ncols)
1178 [ # # ]: 0 : ereport(ERROR,
1179 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1180 : : errmsg_plural("logical replication column %d not found in tuple: only %d column received",
1181 : : "logical replication column %d not found in tuple: only %d columns received",
1182 : : tupleData->ncols,
1183 : : remoteattnum + 1, tupleData->ncols)));
1184 : :
1185 [ + - ]: 68843 : if (tupleData->colstatus[remoteattnum] != LOGICALREP_COLUMN_UNCHANGED)
1186 : : {
1187 : 68843 : StringInfo colvalue = &tupleData->colvalues[remoteattnum];
1188 : :
1189 : : /* Set attnum for error callback */
1190 : 68843 : remote_ctx.remote_attnum = remoteattnum;
1191 : :
1192 [ + + ]: 68843 : if (tupleData->colstatus[remoteattnum] == LOGICALREP_COLUMN_TEXT)
1193 : : {
1194 : : Oid typinput;
1195 : : Oid typioparam;
1196 : :
1197 : 25439 : getTypeInputInfo(att->atttypid, &typinput, &typioparam);
1198 : 50878 : slot->tts_values[i] =
1199 : 25439 : OidInputFunctionCall(typinput, colvalue->data,
1200 : : typioparam, att->atttypmod);
1201 : 25439 : slot->tts_isnull[i] = false;
1202 : : }
1203 [ + + ]: 43404 : else if (tupleData->colstatus[remoteattnum] == LOGICALREP_COLUMN_BINARY)
1204 : : {
1205 : : Oid typreceive;
1206 : : Oid typioparam;
1207 : :
1208 : : /*
1209 : : * In some code paths we may be asked to re-parse the same
1210 : : * tuple data. Reset the StringInfo's cursor so that works.
1211 : : */
1212 : 43356 : colvalue->cursor = 0;
1213 : :
1214 : 43356 : getTypeBinaryInputInfo(att->atttypid, &typreceive, &typioparam);
1215 : 86712 : slot->tts_values[i] =
1216 : 43356 : OidReceiveFunctionCall(typreceive, colvalue,
1217 : : typioparam, att->atttypmod);
1218 : :
1219 : : /* Trouble if it didn't eat the whole buffer */
1220 [ - + ]: 43356 : if (colvalue->cursor != colvalue->len)
1221 [ # # ]: 0 : ereport(ERROR,
1222 : : (errcode(ERRCODE_INVALID_BINARY_REPRESENTATION),
1223 : : errmsg("incorrect binary data format in logical replication column %d",
1224 : : remoteattnum + 1)));
1225 : 43356 : slot->tts_isnull[i] = false;
1226 : : }
1227 : : else
1228 : : {
1229 : : /* must be LOGICALREP_COLUMN_NULL */
1230 : 48 : slot->tts_values[i] = (Datum) 0;
1231 : 48 : slot->tts_isnull[i] = true;
1232 : : }
1233 : :
1234 : : /* Reset attnum for error callback */
1235 : 68843 : remote_ctx.remote_attnum = -1;
1236 : : }
1237 : : }
1238 : :
1239 : : /* And finally, declare that "slot" contains a valid virtual tuple */
1240 : 31927 : ExecStoreVirtualTuple(slot);
1241 : 31927 : }
1242 : :
1243 : : /*
1244 : : * Handle BEGIN message.
1245 : : */
1246 : : static void
1247 : 549 : apply_handle_begin(StringInfo s)
1248 : : {
1249 : : LogicalRepBeginData begin_data;
1250 : :
1251 : : /* There must not be an active streaming transaction. */
1252 : : Assert(!TransactionIdIsValid(stream_xid));
1253 : :
1254 : 549 : logicalrep_read_begin(s, &begin_data);
1255 : 549 : set_remote_transaction_info(begin_data.xid, begin_data.final_lsn);
1256 : :
1257 : 549 : maybe_start_skipping_changes(begin_data.final_lsn);
1258 : :
1259 : 549 : in_remote_transaction = true;
1260 : :
1261 : 549 : pgstat_report_activity(STATE_RUNNING, NULL);
1262 : 549 : }
1263 : :
1264 : : /*
1265 : : * Handle COMMIT message.
1266 : : *
1267 : : * TODO, support tracking of multiple origins
1268 : : */
1269 : : static void
1270 : 459 : apply_handle_commit(StringInfo s)
1271 : : {
1272 : : LogicalRepCommitData commit_data;
1273 : :
1274 : 459 : logicalrep_read_commit(s, &commit_data);
1275 : :
1276 [ - + ]: 459 : if (commit_data.commit_lsn != remote_ctx.finish_lsn)
1277 [ # # ]: 0 : ereport(ERROR,
1278 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1279 : : errmsg_internal("incorrect commit LSN %X/%08X in commit message (expected %X/%08X)",
1280 : : LSN_FORMAT_ARGS(commit_data.commit_lsn),
1281 : : LSN_FORMAT_ARGS(remote_ctx.finish_lsn))));
1282 : :
1283 : 459 : apply_handle_commit_internal(&commit_data);
1284 : :
1285 : : /*
1286 : : * Process any tables that are being synchronized in parallel, as well as
1287 : : * any newly added tables or sequences.
1288 : : */
1289 : 459 : ProcessSyncingRelations(commit_data.end_lsn);
1290 : :
1291 : 459 : pgstat_report_activity(STATE_IDLE, NULL);
1292 : 459 : reset_apply_remote_context();
1293 : 459 : }
1294 : :
1295 : : /*
1296 : : * Handle BEGIN PREPARE message.
1297 : : */
1298 : : static void
1299 : 17 : apply_handle_begin_prepare(StringInfo s)
1300 : : {
1301 : : LogicalRepPreparedTxnData begin_data;
1302 : :
1303 : : /* Tablesync should never receive prepare. */
1304 [ - + ]: 17 : if (am_tablesync_worker())
1305 [ # # ]: 0 : ereport(ERROR,
1306 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1307 : : errmsg_internal("tablesync worker received a BEGIN PREPARE message")));
1308 : :
1309 : : /* There must not be an active streaming transaction. */
1310 : : Assert(!TransactionIdIsValid(stream_xid));
1311 : :
1312 : 17 : logicalrep_read_begin_prepare(s, &begin_data);
1313 : 17 : set_remote_transaction_info(begin_data.xid, begin_data.prepare_lsn);
1314 : :
1315 : 17 : maybe_start_skipping_changes(begin_data.prepare_lsn);
1316 : :
1317 : 17 : in_remote_transaction = true;
1318 : :
1319 : 17 : pgstat_report_activity(STATE_RUNNING, NULL);
1320 : 17 : }
1321 : :
1322 : : /*
1323 : : * Common function to prepare the GID.
1324 : : */
1325 : : static void
1326 : 26 : apply_handle_prepare_internal(LogicalRepPreparedTxnData *prepare_data)
1327 : : {
1328 : : char gid[GIDSIZE];
1329 : :
1330 : : /*
1331 : : * Compute unique GID for two_phase transactions. We don't use GID of
1332 : : * prepared transaction sent by server as that can lead to deadlock when
1333 : : * we have multiple subscriptions from same node point to publications on
1334 : : * the same node. See comments atop worker.c
1335 : : */
1336 : 26 : TwoPhaseTransactionGid(MySubscription->oid, prepare_data->xid,
1337 : : gid, sizeof(gid));
1338 : :
1339 : : /*
1340 : : * BeginTransactionBlock is necessary to balance the EndTransactionBlock
1341 : : * called within the PrepareTransactionBlock below.
1342 : : */
1343 [ + - ]: 26 : if (!IsTransactionBlock())
1344 : : {
1345 : 26 : BeginTransactionBlock();
1346 : 26 : CommitTransactionCommand(); /* Completes the preceding Begin command. */
1347 : : }
1348 : :
1349 : : /*
1350 : : * Update origin state so we can restart streaming from correct position
1351 : : * in case of crash.
1352 : : */
1353 : 26 : replorigin_xact_state.origin_lsn = prepare_data->end_lsn;
1354 : 26 : replorigin_xact_state.origin_timestamp = prepare_data->prepare_time;
1355 : :
1356 : 26 : PrepareTransactionBlock(gid);
1357 : 26 : }
1358 : :
1359 : : /*
1360 : : * Handle PREPARE message.
1361 : : */
1362 : : static void
1363 : 16 : apply_handle_prepare(StringInfo s)
1364 : : {
1365 : : LogicalRepPreparedTxnData prepare_data;
1366 : :
1367 : 16 : logicalrep_read_prepare(s, &prepare_data);
1368 : :
1369 [ - + ]: 16 : if (prepare_data.prepare_lsn != remote_ctx.finish_lsn)
1370 [ # # ]: 0 : ereport(ERROR,
1371 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1372 : : errmsg_internal("incorrect prepare LSN %X/%08X in prepare message (expected %X/%08X)",
1373 : : LSN_FORMAT_ARGS(prepare_data.prepare_lsn),
1374 : : LSN_FORMAT_ARGS(remote_ctx.finish_lsn))));
1375 : :
1376 : : /*
1377 : : * Unlike commit, here, we always prepare the transaction even though no
1378 : : * change has happened in this transaction or all changes are skipped. It
1379 : : * is done this way because at commit prepared time, we won't know whether
1380 : : * we have skipped preparing a transaction because of those reasons.
1381 : : *
1382 : : * XXX, We can optimize such that at commit prepared time, we first check
1383 : : * whether we have prepared the transaction or not but that doesn't seem
1384 : : * worthwhile because such cases shouldn't be common.
1385 : : */
1386 : 16 : begin_replication_step();
1387 : :
1388 : 16 : apply_handle_prepare_internal(&prepare_data);
1389 : :
1390 : 16 : end_replication_step();
1391 : 16 : CommitTransactionCommand();
1392 : 15 : pgstat_report_stat(false);
1393 : :
1394 : : /*
1395 : : * It is okay not to set the local_end LSN for the prepare because we
1396 : : * always flush the prepare record. So, we can send the acknowledgment of
1397 : : * the remote_end LSN as soon as prepare is finished.
1398 : : *
1399 : : * XXX For the sake of consistency with commit, we could have set it with
1400 : : * the LSN of prepare but as of now we don't track that value similar to
1401 : : * XactLastCommitEnd, and adding it for this purpose doesn't seems worth
1402 : : * it.
1403 : : */
1404 : 15 : store_flush_position(prepare_data.end_lsn, InvalidXLogRecPtr);
1405 : :
1406 : 15 : in_remote_transaction = false;
1407 : :
1408 : : /*
1409 : : * Process any tables that are being synchronized in parallel, as well as
1410 : : * any newly added tables or sequences.
1411 : : */
1412 : 15 : ProcessSyncingRelations(prepare_data.end_lsn);
1413 : :
1414 : : /*
1415 : : * Since we have already prepared the transaction, in a case where the
1416 : : * server crashes before clearing the subskiplsn, it will be left but the
1417 : : * transaction won't be resent. But that's okay because it's a rare case
1418 : : * and the subskiplsn will be cleared when finishing the next transaction.
1419 : : */
1420 : 15 : stop_skipping_changes();
1421 : 15 : clear_subscription_skip_lsn(prepare_data.prepare_lsn);
1422 : :
1423 : 15 : pgstat_report_activity(STATE_IDLE, NULL);
1424 : 15 : reset_apply_remote_context();
1425 : 15 : }
1426 : :
1427 : : /*
1428 : : * Handle a COMMIT PREPARED of a previously PREPARED transaction.
1429 : : *
1430 : : * Note that we don't need to wait here if the transaction was prepared in a
1431 : : * parallel apply worker. In that case, we have already waited for the prepare
1432 : : * to finish in apply_handle_stream_prepare() which will ensure all the
1433 : : * operations in that transaction have happened in the subscriber, so no
1434 : : * concurrent transaction can cause deadlock or transaction dependency issues.
1435 : : */
1436 : : static void
1437 : 22 : apply_handle_commit_prepared(StringInfo s)
1438 : : {
1439 : : LogicalRepCommitPreparedTxnData prepare_data;
1440 : : char gid[GIDSIZE];
1441 : :
1442 : 22 : logicalrep_read_commit_prepared(s, &prepare_data);
1443 : 22 : set_remote_transaction_info(prepare_data.xid, prepare_data.commit_lsn);
1444 : :
1445 : : /* Compute GID for two_phase transactions. */
1446 : 22 : TwoPhaseTransactionGid(MySubscription->oid, prepare_data.xid,
1447 : : gid, sizeof(gid));
1448 : :
1449 : : /* There is no transaction when COMMIT PREPARED is called */
1450 : 22 : begin_replication_step();
1451 : :
1452 : : /*
1453 : : * Update origin state so we can restart streaming from correct position
1454 : : * in case of crash.
1455 : : */
1456 : 22 : replorigin_xact_state.origin_lsn = prepare_data.end_lsn;
1457 : 22 : replorigin_xact_state.origin_timestamp = prepare_data.commit_time;
1458 : :
1459 : 22 : FinishPreparedTransaction(gid, true);
1460 : 22 : end_replication_step();
1461 : 22 : CommitTransactionCommand();
1462 : 22 : pgstat_report_stat(false);
1463 : :
1464 : 22 : store_flush_position(prepare_data.end_lsn, XactLastCommitEnd);
1465 : 22 : in_remote_transaction = false;
1466 : :
1467 : : /*
1468 : : * Process any tables that are being synchronized in parallel, as well as
1469 : : * any newly added tables or sequences.
1470 : : */
1471 : 22 : ProcessSyncingRelations(prepare_data.end_lsn);
1472 : :
1473 : 22 : clear_subscription_skip_lsn(prepare_data.end_lsn);
1474 : :
1475 : 22 : pgstat_report_activity(STATE_IDLE, NULL);
1476 : 22 : reset_apply_remote_context();
1477 : 22 : }
1478 : :
1479 : : /*
1480 : : * Handle a ROLLBACK PREPARED of a previously PREPARED TRANSACTION.
1481 : : *
1482 : : * Note that we don't need to wait here if the transaction was prepared in a
1483 : : * parallel apply worker. In that case, we have already waited for the prepare
1484 : : * to finish in apply_handle_stream_prepare() which will ensure all the
1485 : : * operations in that transaction have happened in the subscriber, so no
1486 : : * concurrent transaction can cause deadlock or transaction dependency issues.
1487 : : */
1488 : : static void
1489 : 5 : apply_handle_rollback_prepared(StringInfo s)
1490 : : {
1491 : : LogicalRepRollbackPreparedTxnData rollback_data;
1492 : : char gid[GIDSIZE];
1493 : :
1494 : 5 : logicalrep_read_rollback_prepared(s, &rollback_data);
1495 : 5 : set_remote_transaction_info(rollback_data.xid, rollback_data.rollback_end_lsn);
1496 : :
1497 : : /* Compute GID for two_phase transactions. */
1498 : 5 : TwoPhaseTransactionGid(MySubscription->oid, rollback_data.xid,
1499 : : gid, sizeof(gid));
1500 : :
1501 : : /*
1502 : : * It is possible that we haven't received prepare because it occurred
1503 : : * before walsender reached a consistent point or the two_phase was still
1504 : : * not enabled by that time, so in such cases, we need to skip rollback
1505 : : * prepared.
1506 : : */
1507 [ + - ]: 5 : if (LookupGXact(gid, rollback_data.prepare_end_lsn,
1508 : : rollback_data.prepare_time))
1509 : : {
1510 : : /*
1511 : : * Update origin state so we can restart streaming from correct
1512 : : * position in case of crash.
1513 : : */
1514 : 5 : replorigin_xact_state.origin_lsn = rollback_data.rollback_end_lsn;
1515 : 5 : replorigin_xact_state.origin_timestamp = rollback_data.rollback_time;
1516 : :
1517 : : /* There is no transaction when ABORT/ROLLBACK PREPARED is called */
1518 : 5 : begin_replication_step();
1519 : 5 : FinishPreparedTransaction(gid, false);
1520 : 5 : end_replication_step();
1521 : 5 : CommitTransactionCommand();
1522 : :
1523 : 5 : clear_subscription_skip_lsn(rollback_data.rollback_end_lsn);
1524 : : }
1525 : :
1526 : 5 : pgstat_report_stat(false);
1527 : :
1528 : : /*
1529 : : * It is okay not to set the local_end LSN for the rollback of prepared
1530 : : * transaction because we always flush the WAL record for it. See
1531 : : * apply_handle_prepare.
1532 : : */
1533 : 5 : store_flush_position(rollback_data.rollback_end_lsn, InvalidXLogRecPtr);
1534 : 5 : in_remote_transaction = false;
1535 : :
1536 : : /*
1537 : : * Process any tables that are being synchronized in parallel, as well as
1538 : : * any newly added tables or sequences.
1539 : : */
1540 : 5 : ProcessSyncingRelations(rollback_data.rollback_end_lsn);
1541 : :
1542 : 5 : pgstat_report_activity(STATE_IDLE, NULL);
1543 : 5 : reset_apply_remote_context();
1544 : 5 : }
1545 : :
1546 : : /*
1547 : : * Handle STREAM PREPARE.
1548 : : */
1549 : : static void
1550 : 15 : apply_handle_stream_prepare(StringInfo s)
1551 : : {
1552 : : LogicalRepPreparedTxnData prepare_data;
1553 : : ParallelApplyWorkerInfo *winfo;
1554 : : TransApplyAction apply_action;
1555 : :
1556 : : /* Save the message before it is consumed. */
1557 : 15 : StringInfoData original_msg = *s;
1558 : :
1559 [ - + ]: 15 : if (in_streamed_transaction)
1560 [ # # ]: 0 : ereport(ERROR,
1561 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1562 : : errmsg_internal("STREAM PREPARE message without STREAM STOP")));
1563 : :
1564 : : /* Tablesync should never receive prepare. */
1565 [ - + ]: 15 : if (am_tablesync_worker())
1566 [ # # ]: 0 : ereport(ERROR,
1567 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1568 : : errmsg_internal("tablesync worker received a STREAM PREPARE message")));
1569 : :
1570 : 15 : logicalrep_read_stream_prepare(s, &prepare_data);
1571 : 15 : set_remote_transaction_info(prepare_data.xid, prepare_data.prepare_lsn);
1572 : :
1573 : 15 : apply_action = get_transaction_apply_action(prepare_data.xid, &winfo);
1574 : :
1575 [ + + + + : 15 : switch (apply_action)
- ]
1576 : : {
1577 : 5 : case TRANS_LEADER_APPLY:
1578 : :
1579 : : /*
1580 : : * The transaction has been serialized to file, so replay all the
1581 : : * spooled operations.
1582 : : */
1583 : 5 : apply_spooled_messages(MyLogicalRepWorker->stream_fileset,
1584 : : prepare_data.xid, prepare_data.prepare_lsn);
1585 : :
1586 : : /* Mark the transaction as prepared. */
1587 : 5 : apply_handle_prepare_internal(&prepare_data);
1588 : :
1589 : 5 : CommitTransactionCommand();
1590 : :
1591 : : /*
1592 : : * It is okay not to set the local_end LSN for the prepare because
1593 : : * we always flush the prepare record. See apply_handle_prepare.
1594 : : */
1595 : 5 : store_flush_position(prepare_data.end_lsn, InvalidXLogRecPtr);
1596 : :
1597 : 5 : in_remote_transaction = false;
1598 : :
1599 : : /* Unlink the files with serialized changes and subxact info. */
1600 : 5 : stream_cleanup_files(MyLogicalRepWorker->subid, prepare_data.xid);
1601 : :
1602 [ - + ]: 5 : elog(DEBUG1, "finished processing the STREAM PREPARE command");
1603 : 5 : break;
1604 : :
1605 : 4 : case TRANS_LEADER_SEND_TO_PARALLEL:
1606 : : Assert(winfo);
1607 : :
1608 [ + - ]: 4 : if (pa_send_data(winfo, s->len, s->data))
1609 : : {
1610 : : /* Finish processing the streaming transaction. */
1611 : 4 : pa_xact_finish(winfo, prepare_data.end_lsn);
1612 : 3 : break;
1613 : : }
1614 : :
1615 : : /*
1616 : : * Switch to serialize mode when we are not able to send the
1617 : : * change to parallel apply worker.
1618 : : */
1619 : 0 : pa_switch_to_partial_serialize(winfo, true);
1620 : :
1621 : : pg_fallthrough;
1622 : 1 : case TRANS_LEADER_PARTIAL_SERIALIZE:
1623 : : Assert(winfo);
1624 : :
1625 : 1 : stream_open_and_write_change(prepare_data.xid,
1626 : : LOGICAL_REP_MSG_STREAM_PREPARE,
1627 : : &original_msg);
1628 : :
1629 : 1 : pa_set_fileset_state(winfo->shared, FS_SERIALIZE_DONE);
1630 : :
1631 : : /* Finish processing the streaming transaction. */
1632 : 1 : pa_xact_finish(winfo, prepare_data.end_lsn);
1633 : 1 : break;
1634 : :
1635 : 5 : case TRANS_PARALLEL_APPLY:
1636 : :
1637 : : /*
1638 : : * If the parallel apply worker is applying spooled messages then
1639 : : * close the file before preparing.
1640 : : */
1641 [ + + ]: 5 : if (stream_fd)
1642 : 1 : stream_close_file();
1643 : :
1644 : 5 : begin_replication_step();
1645 : :
1646 : : /* Mark the transaction as prepared. */
1647 : 5 : apply_handle_prepare_internal(&prepare_data);
1648 : :
1649 : 5 : end_replication_step();
1650 : :
1651 : 5 : CommitTransactionCommand();
1652 : :
1653 : : /*
1654 : : * It is okay not to set the local_end LSN for the prepare because
1655 : : * we always flush the prepare record. See apply_handle_prepare.
1656 : : */
1657 : 4 : MyParallelShared->last_commit_end = InvalidXLogRecPtr;
1658 : :
1659 : 4 : pa_set_xact_state(MyParallelShared, PARALLEL_TRANS_FINISHED);
1660 : 4 : pa_unlock_transaction(MyParallelShared->xid, AccessExclusiveLock);
1661 : :
1662 : 4 : pa_reset_subtrans();
1663 : :
1664 [ + + ]: 4 : elog(DEBUG1, "finished processing the STREAM PREPARE command");
1665 : 4 : break;
1666 : :
1667 : 0 : default:
1668 [ # # ]: 0 : elog(ERROR, "unexpected apply action: %d", (int) apply_action);
1669 : : break;
1670 : : }
1671 : :
1672 : 13 : pgstat_report_stat(false);
1673 : :
1674 : : /*
1675 : : * Process any tables that are being synchronized in parallel, as well as
1676 : : * any newly added tables or sequences.
1677 : : */
1678 : 13 : ProcessSyncingRelations(prepare_data.end_lsn);
1679 : :
1680 : : /*
1681 : : * Similar to prepare case, the subskiplsn could be left in a case of
1682 : : * server crash but it's okay. See the comments in apply_handle_prepare().
1683 : : */
1684 : 13 : stop_skipping_changes();
1685 : 13 : clear_subscription_skip_lsn(prepare_data.prepare_lsn);
1686 : :
1687 : 13 : pgstat_report_activity(STATE_IDLE, NULL);
1688 : :
1689 : 13 : reset_apply_remote_context();
1690 : 13 : }
1691 : :
1692 : : /*
1693 : : * Handle ORIGIN message.
1694 : : *
1695 : : * TODO, support tracking of multiple origins
1696 : : */
1697 : : static void
1698 : 7 : apply_handle_origin(StringInfo s)
1699 : : {
1700 : : /*
1701 : : * ORIGIN message can only come inside streaming transaction or inside
1702 : : * remote transaction and before any actual writes.
1703 : : */
1704 [ + + ]: 7 : if (!in_streamed_transaction &&
1705 [ + - - + ]: 10 : (!in_remote_transaction ||
1706 [ - - ]: 5 : (IsTransactionState() && !am_tablesync_worker())))
1707 [ # # ]: 0 : ereport(ERROR,
1708 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1709 : : errmsg_internal("ORIGIN message sent out of order")));
1710 : 7 : }
1711 : :
1712 : : /*
1713 : : * Initialize fileset (if not already done).
1714 : : *
1715 : : * Create a new file when first_segment is true, otherwise open the existing
1716 : : * file.
1717 : : */
1718 : : void
1719 : 362 : stream_start_internal(TransactionId xid, bool first_segment)
1720 : : {
1721 : 362 : begin_replication_step();
1722 : :
1723 : : /*
1724 : : * Initialize the worker's stream_fileset if we haven't yet. This will be
1725 : : * used for the entire duration of the worker so create it in a permanent
1726 : : * context. We create this on the very first streaming message from any
1727 : : * transaction and then use it for this and other streaming transactions.
1728 : : * Now, we could create a fileset at the start of the worker as well but
1729 : : * then we won't be sure that it will ever be used.
1730 : : */
1731 [ + + ]: 362 : if (!MyLogicalRepWorker->stream_fileset)
1732 : : {
1733 : : MemoryContext oldctx;
1734 : :
1735 : 14 : oldctx = MemoryContextSwitchTo(ApplyContext);
1736 : :
1737 : 14 : MyLogicalRepWorker->stream_fileset = palloc_object(FileSet);
1738 : 14 : FileSetInit(MyLogicalRepWorker->stream_fileset);
1739 : :
1740 : 14 : MemoryContextSwitchTo(oldctx);
1741 : : }
1742 : :
1743 : : /* Open the spool file for this transaction. */
1744 : 362 : stream_open_file(MyLogicalRepWorker->subid, xid, first_segment);
1745 : :
1746 : : /* If this is not the first segment, open existing subxact file. */
1747 [ + + ]: 362 : if (!first_segment)
1748 : 330 : subxact_info_read(MyLogicalRepWorker->subid, xid);
1749 : :
1750 : 362 : end_replication_step();
1751 : 362 : }
1752 : :
1753 : : /*
1754 : : * Handle STREAM START message.
1755 : : */
1756 : : static void
1757 : 842 : apply_handle_stream_start(StringInfo s)
1758 : : {
1759 : : bool first_segment;
1760 : : ParallelApplyWorkerInfo *winfo;
1761 : : TransApplyAction apply_action;
1762 : :
1763 : : /* Save the message before it is consumed. */
1764 : 842 : StringInfoData original_msg = *s;
1765 : :
1766 [ - + ]: 842 : if (in_streamed_transaction)
1767 [ # # ]: 0 : ereport(ERROR,
1768 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1769 : : errmsg_internal("duplicate STREAM START message")));
1770 : :
1771 : : /* There must not be an active streaming transaction. */
1772 : : Assert(!TransactionIdIsValid(stream_xid));
1773 : :
1774 : : /* notify handle methods we're processing a remote transaction */
1775 : 842 : in_streamed_transaction = true;
1776 : :
1777 : : /* extract XID of the top-level transaction */
1778 : 842 : stream_xid = logicalrep_read_stream_start(s, &first_segment);
1779 : :
1780 [ - + ]: 842 : if (!TransactionIdIsValid(stream_xid))
1781 [ # # ]: 0 : ereport(ERROR,
1782 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1783 : : errmsg_internal("invalid transaction ID in streamed replication transaction")));
1784 : :
1785 : : /*
1786 : : * The final LSN of the streamed transaction is known only when its commit
1787 : : * record arrives.
1788 : : */
1789 : 842 : set_remote_transaction_info(stream_xid, InvalidXLogRecPtr);
1790 : :
1791 : : /* Try to allocate a worker for the streaming transaction. */
1792 [ + + ]: 842 : if (first_segment)
1793 : 86 : pa_allocate_worker(stream_xid);
1794 : :
1795 : 842 : apply_action = get_transaction_apply_action(stream_xid, &winfo);
1796 : :
1797 [ + + + + : 842 : switch (apply_action)
- ]
1798 : : {
1799 : 342 : case TRANS_LEADER_SERIALIZE:
1800 : :
1801 : : /*
1802 : : * Function stream_start_internal starts a transaction. This
1803 : : * transaction will be committed on the stream stop unless it is a
1804 : : * tablesync worker in which case it will be committed after
1805 : : * processing all the messages. We need this transaction for
1806 : : * handling the BufFile, used for serializing the streaming data
1807 : : * and subxact info.
1808 : : */
1809 : 342 : stream_start_internal(stream_xid, first_segment);
1810 : 342 : break;
1811 : :
1812 : 244 : case TRANS_LEADER_SEND_TO_PARALLEL:
1813 : : Assert(winfo);
1814 : :
1815 : : /*
1816 : : * Once we start serializing the changes, the parallel apply
1817 : : * worker will wait for the leader to release the stream lock
1818 : : * until the end of the transaction. So, we don't need to release
1819 : : * the lock or increment the stream count in that case.
1820 : : */
1821 [ + + ]: 244 : if (pa_send_data(winfo, s->len, s->data))
1822 : : {
1823 : : /*
1824 : : * Unlock the shared object lock so that the parallel apply
1825 : : * worker can continue to receive changes.
1826 : : */
1827 [ + + ]: 240 : if (!first_segment)
1828 : 215 : pa_unlock_stream(winfo->shared->xid, AccessExclusiveLock);
1829 : :
1830 : : /*
1831 : : * Increment the number of streaming blocks waiting to be
1832 : : * processed by parallel apply worker.
1833 : : */
1834 : 240 : pg_atomic_add_fetch_u32(&winfo->shared->pending_stream_count, 1);
1835 : :
1836 : : /* Cache the parallel apply worker for this transaction. */
1837 : 240 : pa_set_stream_apply_worker(winfo);
1838 : 240 : break;
1839 : : }
1840 : :
1841 : : /*
1842 : : * Switch to serialize mode when we are not able to send the
1843 : : * change to parallel apply worker.
1844 : : */
1845 : 4 : pa_switch_to_partial_serialize(winfo, !first_segment);
1846 : :
1847 : : pg_fallthrough;
1848 : 15 : case TRANS_LEADER_PARTIAL_SERIALIZE:
1849 : : Assert(winfo);
1850 : :
1851 : : /*
1852 : : * Open the spool file unless it was already opened when switching
1853 : : * to serialize mode. The transaction started in
1854 : : * stream_start_internal will be committed on the stream stop.
1855 : : */
1856 [ + + ]: 15 : if (apply_action != TRANS_LEADER_SEND_TO_PARALLEL)
1857 : 11 : stream_start_internal(stream_xid, first_segment);
1858 : :
1859 : 15 : stream_write_change(LOGICAL_REP_MSG_STREAM_START, &original_msg);
1860 : :
1861 : : /* Cache the parallel apply worker for this transaction. */
1862 : 15 : pa_set_stream_apply_worker(winfo);
1863 : 15 : break;
1864 : :
1865 : 245 : case TRANS_PARALLEL_APPLY:
1866 [ + + ]: 245 : if (first_segment)
1867 : : {
1868 : : /* Hold the lock until the end of the transaction. */
1869 : 29 : pa_lock_transaction(MyParallelShared->xid, AccessExclusiveLock);
1870 : 29 : pa_set_xact_state(MyParallelShared, PARALLEL_TRANS_STARTED);
1871 : :
1872 : : /*
1873 : : * Signal the leader apply worker, as it may be waiting for
1874 : : * us.
1875 : : */
1876 : 29 : logicalrep_worker_wakeup(WORKERTYPE_APPLY,
1877 : 29 : MyLogicalRepWorker->subid, InvalidOid);
1878 : : }
1879 : :
1880 : 245 : parallel_stream_nchanges = 0;
1881 : 245 : break;
1882 : :
1883 : 0 : default:
1884 [ # # ]: 0 : elog(ERROR, "unexpected apply action: %d", (int) apply_action);
1885 : : break;
1886 : : }
1887 : :
1888 : 842 : pgstat_report_activity(STATE_RUNNING, NULL);
1889 : 842 : }
1890 : :
1891 : : /*
1892 : : * Update the information about subxacts and close the file.
1893 : : *
1894 : : * This function should be called when the stream_start_internal function has
1895 : : * been called.
1896 : : */
1897 : : void
1898 : 362 : stream_stop_internal(TransactionId xid)
1899 : : {
1900 : : /*
1901 : : * Serialize information about subxacts for the toplevel transaction, then
1902 : : * close the stream messages spool file.
1903 : : */
1904 : 362 : subxact_info_write(MyLogicalRepWorker->subid, xid);
1905 : 362 : stream_close_file();
1906 : :
1907 : : /* We must be in a valid transaction state */
1908 : : Assert(IsTransactionState());
1909 : :
1910 : : /* Commit the per-stream transaction */
1911 : 362 : CommitTransactionCommand();
1912 : :
1913 : : /* Reset per-stream context */
1914 : 362 : MemoryContextReset(LogicalStreamingContext);
1915 : 362 : }
1916 : :
1917 : : /*
1918 : : * Handle STREAM STOP message.
1919 : : */
1920 : : static void
1921 : 841 : apply_handle_stream_stop(StringInfo s)
1922 : : {
1923 : : ParallelApplyWorkerInfo *winfo;
1924 : : TransApplyAction apply_action;
1925 : :
1926 [ - + ]: 841 : if (!in_streamed_transaction)
1927 [ # # ]: 0 : ereport(ERROR,
1928 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
1929 : : errmsg_internal("STREAM STOP message without STREAM START")));
1930 : :
1931 : 841 : apply_action = get_transaction_apply_action(stream_xid, &winfo);
1932 : :
1933 [ + + + + : 841 : switch (apply_action)
- ]
1934 : : {
1935 : 342 : case TRANS_LEADER_SERIALIZE:
1936 : 342 : stream_stop_internal(stream_xid);
1937 : 342 : break;
1938 : :
1939 : 240 : case TRANS_LEADER_SEND_TO_PARALLEL:
1940 : : Assert(winfo);
1941 : :
1942 : : /*
1943 : : * Lock before sending the STREAM_STOP message so that the leader
1944 : : * can hold the lock first and the parallel apply worker will wait
1945 : : * for leader to release the lock. See Locking Considerations atop
1946 : : * applyparallelworker.c.
1947 : : */
1948 : 240 : pa_lock_stream(winfo->shared->xid, AccessExclusiveLock);
1949 : :
1950 [ + - ]: 240 : if (pa_send_data(winfo, s->len, s->data))
1951 : : {
1952 : 240 : pa_set_stream_apply_worker(NULL);
1953 : 240 : break;
1954 : : }
1955 : :
1956 : : /*
1957 : : * Switch to serialize mode when we are not able to send the
1958 : : * change to parallel apply worker.
1959 : : */
1960 : 0 : pa_switch_to_partial_serialize(winfo, true);
1961 : :
1962 : : pg_fallthrough;
1963 : 15 : case TRANS_LEADER_PARTIAL_SERIALIZE:
1964 : 15 : stream_write_change(LOGICAL_REP_MSG_STREAM_STOP, s);
1965 : 15 : stream_stop_internal(stream_xid);
1966 : 15 : pa_set_stream_apply_worker(NULL);
1967 : 15 : break;
1968 : :
1969 : 244 : case TRANS_PARALLEL_APPLY:
1970 [ + + ]: 244 : elog(DEBUG1, "applied %u changes in the streaming chunk",
1971 : : parallel_stream_nchanges);
1972 : :
1973 : : /*
1974 : : * By the time parallel apply worker is processing the changes in
1975 : : * the current streaming block, the leader apply worker may have
1976 : : * sent multiple streaming blocks. This can lead to parallel apply
1977 : : * worker start waiting even when there are more chunk of streams
1978 : : * in the queue. So, try to lock only if there is no message left
1979 : : * in the queue. See Locking Considerations atop
1980 : : * applyparallelworker.c.
1981 : : *
1982 : : * Note that here we have a race condition where we can start
1983 : : * waiting even when there are pending streaming chunks. This can
1984 : : * happen if the leader sends another streaming block and acquires
1985 : : * the stream lock again after the parallel apply worker checks
1986 : : * that there is no pending streaming block and before it actually
1987 : : * starts waiting on a lock. We can handle this case by not
1988 : : * allowing the leader to increment the stream block count during
1989 : : * the time parallel apply worker acquires the lock but it is not
1990 : : * clear whether that is worth the complexity.
1991 : : *
1992 : : * Now, if this missed chunk contains rollback to savepoint, then
1993 : : * there is a risk of deadlock which probably shouldn't happen
1994 : : * after restart.
1995 : : */
1996 : 244 : pa_decr_and_wait_stream_block();
1997 : 242 : break;
1998 : :
1999 : 0 : default:
2000 [ # # ]: 0 : elog(ERROR, "unexpected apply action: %d", (int) apply_action);
2001 : : break;
2002 : : }
2003 : :
2004 : 839 : in_streamed_transaction = false;
2005 : 839 : stream_xid = InvalidTransactionId;
2006 : :
2007 : : /*
2008 : : * The parallel apply worker could be in a transaction in which case we
2009 : : * need to report the state as STATE_IDLEINTRANSACTION.
2010 : : */
2011 [ + + ]: 839 : if (IsTransactionOrTransactionBlock())
2012 : 242 : pgstat_report_activity(STATE_IDLEINTRANSACTION, NULL);
2013 : : else
2014 : 597 : pgstat_report_activity(STATE_IDLE, NULL);
2015 : :
2016 : 839 : reset_apply_remote_context();
2017 : 839 : }
2018 : :
2019 : : /*
2020 : : * Helper function to handle STREAM ABORT message when the transaction was
2021 : : * serialized to file.
2022 : : */
2023 : : static void
2024 : 14 : stream_abort_internal(TransactionId xid, TransactionId subxid)
2025 : : {
2026 : : /*
2027 : : * If the two XIDs are the same, it's in fact abort of toplevel xact, so
2028 : : * just delete the files with serialized info.
2029 : : */
2030 [ + + ]: 14 : if (xid == subxid)
2031 : 1 : stream_cleanup_files(MyLogicalRepWorker->subid, xid);
2032 : : else
2033 : : {
2034 : : /*
2035 : : * OK, so it's a subxact. We need to read the subxact file for the
2036 : : * toplevel transaction, determine the offset tracked for the subxact,
2037 : : * and truncate the file with changes. We also remove the subxacts
2038 : : * with higher offsets (or rather higher XIDs).
2039 : : *
2040 : : * We intentionally scan the array from the tail, because we're likely
2041 : : * aborting a change for the most recent subtransactions.
2042 : : *
2043 : : * We can't use the binary search here as subxact XIDs won't
2044 : : * necessarily arrive in sorted order, consider the case where we have
2045 : : * released the savepoint for multiple subtransactions and then
2046 : : * performed rollback to savepoint for one of the earlier
2047 : : * sub-transaction.
2048 : : */
2049 : : int64 i;
2050 : : int64 subidx;
2051 : : BufFile *fd;
2052 : 13 : bool found = false;
2053 : : char path[MAXPGPATH];
2054 : :
2055 : 13 : subidx = -1;
2056 : 13 : begin_replication_step();
2057 : 13 : subxact_info_read(MyLogicalRepWorker->subid, xid);
2058 : :
2059 [ + + ]: 15 : for (i = subxact_data.nsubxacts; i > 0; i--)
2060 : : {
2061 [ + + ]: 11 : if (subxact_data.subxacts[i - 1].xid == subxid)
2062 : : {
2063 : 9 : subidx = (i - 1);
2064 : 9 : found = true;
2065 : 9 : break;
2066 : : }
2067 : : }
2068 : :
2069 : : /*
2070 : : * If it's an empty sub-transaction then we will not find the subxid
2071 : : * here so just cleanup the subxact info and return.
2072 : : */
2073 [ + + ]: 13 : if (!found)
2074 : : {
2075 : : /* Cleanup the subxact info */
2076 : 4 : cleanup_subxact_info();
2077 : 4 : end_replication_step();
2078 : 4 : CommitTransactionCommand();
2079 : 4 : return;
2080 : : }
2081 : :
2082 : : /* open the changes file */
2083 : 9 : changes_filename(path, MyLogicalRepWorker->subid, xid);
2084 : 9 : fd = BufFileOpenFileSet(MyLogicalRepWorker->stream_fileset, path,
2085 : : O_RDWR, false);
2086 : :
2087 : : /* OK, truncate the file at the right offset */
2088 : 9 : BufFileTruncateFileSet(fd, subxact_data.subxacts[subidx].fileno,
2089 : 9 : subxact_data.subxacts[subidx].offset);
2090 : 9 : BufFileClose(fd);
2091 : :
2092 : : /* discard the subxacts added later */
2093 : 9 : subxact_data.nsubxacts = subidx;
2094 : :
2095 : : /* write the updated subxact list */
2096 : 9 : subxact_info_write(MyLogicalRepWorker->subid, xid);
2097 : :
2098 : 9 : end_replication_step();
2099 : 9 : CommitTransactionCommand();
2100 : : }
2101 : : }
2102 : :
2103 : : /*
2104 : : * Handle STREAM ABORT message.
2105 : : */
2106 : : static void
2107 : 38 : apply_handle_stream_abort(StringInfo s)
2108 : : {
2109 : : TransactionId xid;
2110 : : TransactionId subxid;
2111 : : LogicalRepStreamAbortData abort_data;
2112 : : ParallelApplyWorkerInfo *winfo;
2113 : : TransApplyAction apply_action;
2114 : :
2115 : : /* Save the message before it is consumed. */
2116 : 38 : StringInfoData original_msg = *s;
2117 : : bool toplevel_xact;
2118 : :
2119 [ - + ]: 38 : if (in_streamed_transaction)
2120 [ # # ]: 0 : ereport(ERROR,
2121 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
2122 : : errmsg_internal("STREAM ABORT message without STREAM STOP")));
2123 : :
2124 : : /* We receive abort information only when we can apply in parallel. */
2125 : 38 : logicalrep_read_stream_abort(s, &abort_data,
2126 : 38 : MyLogicalRepWorker->parallel_apply);
2127 : :
2128 : 38 : xid = abort_data.xid;
2129 : 38 : subxid = abort_data.subxid;
2130 : 38 : toplevel_xact = (xid == subxid);
2131 : :
2132 : : /*
2133 : : * Record the xid of the (sub)transaction being aborted, so that the error
2134 : : * context names whatever failed. Note this is the top-level xid itself
2135 : : * when a top-level transaction aborts, and a subxid only when a
2136 : : * subtransaction rolls back. See set_remote_transaction_info().
2137 : : */
2138 : 38 : set_remote_transaction_info(subxid, abort_data.abort_lsn);
2139 : :
2140 : 38 : apply_action = get_transaction_apply_action(xid, &winfo);
2141 : :
2142 [ + + + + : 38 : switch (apply_action)
- ]
2143 : : {
2144 : 14 : case TRANS_LEADER_APPLY:
2145 : :
2146 : : /*
2147 : : * We are in the leader apply worker and the transaction has been
2148 : : * serialized to file.
2149 : : */
2150 : 14 : stream_abort_internal(xid, subxid);
2151 : :
2152 [ - + ]: 14 : elog(DEBUG1, "finished processing the STREAM ABORT command");
2153 : 14 : break;
2154 : :
2155 : 10 : case TRANS_LEADER_SEND_TO_PARALLEL:
2156 : : Assert(winfo);
2157 : :
2158 : : /*
2159 : : * For the case of aborting the subtransaction, we increment the
2160 : : * number of streaming blocks and take the lock again before
2161 : : * sending the STREAM_ABORT to ensure that the parallel apply
2162 : : * worker will wait on the lock for the next set of changes after
2163 : : * processing the STREAM_ABORT message if it is not already
2164 : : * waiting for STREAM_STOP message.
2165 : : *
2166 : : * It is important to perform this locking before sending the
2167 : : * STREAM_ABORT message so that the leader can hold the lock first
2168 : : * and the parallel apply worker will wait for the leader to
2169 : : * release the lock. This is the same as what we do in
2170 : : * apply_handle_stream_stop. See Locking Considerations atop
2171 : : * applyparallelworker.c.
2172 : : */
2173 [ + + ]: 10 : if (!toplevel_xact)
2174 : : {
2175 : 9 : pa_unlock_stream(xid, AccessExclusiveLock);
2176 : 9 : pg_atomic_add_fetch_u32(&winfo->shared->pending_stream_count, 1);
2177 : 9 : pa_lock_stream(xid, AccessExclusiveLock);
2178 : : }
2179 : :
2180 [ + - ]: 10 : if (pa_send_data(winfo, s->len, s->data))
2181 : : {
2182 : : /*
2183 : : * Unlike STREAM_COMMIT and STREAM_PREPARE, we don't need to
2184 : : * wait here for the parallel apply worker to finish as that
2185 : : * is not required to maintain the commit order and won't have
2186 : : * the risk of failures due to transaction dependencies and
2187 : : * deadlocks. However, it is possible that before the parallel
2188 : : * worker finishes and we clear the worker info, the xid
2189 : : * wraparound happens on the upstream and a new transaction
2190 : : * with the same xid can appear and that can lead to duplicate
2191 : : * entries in ParallelApplyTxnHash. Yet another problem could
2192 : : * be that we may have serialized the changes in partial
2193 : : * serialize mode and the file containing xact changes may
2194 : : * already exist, and after xid wraparound trying to create
2195 : : * the file for the same xid can lead to an error. To avoid
2196 : : * these problems, we decide to wait for the aborts to finish.
2197 : : *
2198 : : * Note, it is okay to not update the flush location position
2199 : : * for aborts as in worst case that means such a transaction
2200 : : * won't be sent again after restart.
2201 : : */
2202 [ + + ]: 10 : if (toplevel_xact)
2203 : 1 : pa_xact_finish(winfo, InvalidXLogRecPtr);
2204 : :
2205 : 10 : break;
2206 : : }
2207 : :
2208 : : /*
2209 : : * Switch to serialize mode when we are not able to send the
2210 : : * change to parallel apply worker.
2211 : : */
2212 : 0 : pa_switch_to_partial_serialize(winfo, true);
2213 : :
2214 : : pg_fallthrough;
2215 : 2 : case TRANS_LEADER_PARTIAL_SERIALIZE:
2216 : : Assert(winfo);
2217 : :
2218 : : /*
2219 : : * Parallel apply worker might have applied some changes, so write
2220 : : * the STREAM_ABORT message so that it can rollback the
2221 : : * subtransaction if needed.
2222 : : */
2223 : 2 : stream_open_and_write_change(xid, LOGICAL_REP_MSG_STREAM_ABORT,
2224 : : &original_msg);
2225 : :
2226 [ + + ]: 2 : if (toplevel_xact)
2227 : : {
2228 : 1 : pa_set_fileset_state(winfo->shared, FS_SERIALIZE_DONE);
2229 : 1 : pa_xact_finish(winfo, InvalidXLogRecPtr);
2230 : : }
2231 : 2 : break;
2232 : :
2233 : 12 : case TRANS_PARALLEL_APPLY:
2234 : :
2235 : : /*
2236 : : * If the parallel apply worker is applying spooled messages then
2237 : : * close the file before aborting.
2238 : : */
2239 [ + + + + ]: 12 : if (toplevel_xact && stream_fd)
2240 : 1 : stream_close_file();
2241 : :
2242 : 12 : pa_stream_abort(&abort_data);
2243 : :
2244 : : /*
2245 : : * We need to wait after processing rollback to savepoint for the
2246 : : * next set of changes.
2247 : : *
2248 : : * We have a race condition here due to which we can start waiting
2249 : : * here when there are more chunk of streams in the queue. See
2250 : : * apply_handle_stream_stop.
2251 : : */
2252 [ + + ]: 12 : if (!toplevel_xact)
2253 : 10 : pa_decr_and_wait_stream_block();
2254 : :
2255 [ + + ]: 12 : elog(DEBUG1, "finished processing the STREAM ABORT command");
2256 : 12 : break;
2257 : :
2258 : 0 : default:
2259 [ # # ]: 0 : elog(ERROR, "unexpected apply action: %d", (int) apply_action);
2260 : : break;
2261 : : }
2262 : :
2263 : 38 : reset_apply_remote_context();
2264 : 38 : }
2265 : :
2266 : : /*
2267 : : * Ensure that the passed location is fileset's end.
2268 : : */
2269 : : static void
2270 : 4 : ensure_last_message(FileSet *stream_fileset, TransactionId xid, int fileno,
2271 : : pgoff_t offset)
2272 : : {
2273 : : char path[MAXPGPATH];
2274 : : BufFile *fd;
2275 : : int last_fileno;
2276 : : pgoff_t last_offset;
2277 : :
2278 : : Assert(!IsTransactionState());
2279 : :
2280 : 4 : begin_replication_step();
2281 : :
2282 : 4 : changes_filename(path, MyLogicalRepWorker->subid, xid);
2283 : :
2284 : 4 : fd = BufFileOpenFileSet(stream_fileset, path, O_RDONLY, false);
2285 : :
2286 : 4 : BufFileSeek(fd, 0, 0, SEEK_END);
2287 : 4 : BufFileTell(fd, &last_fileno, &last_offset);
2288 : :
2289 : 4 : BufFileClose(fd);
2290 : :
2291 : 4 : end_replication_step();
2292 : :
2293 [ + - - + ]: 4 : if (last_fileno != fileno || last_offset != offset)
2294 [ # # ]: 0 : elog(ERROR, "unexpected message left in streaming transaction's changes file \"%s\"",
2295 : : path);
2296 : 4 : }
2297 : :
2298 : : /*
2299 : : * Common spoolfile processing.
2300 : : */
2301 : : void
2302 : 31 : apply_spooled_messages(FileSet *stream_fileset, TransactionId xid,
2303 : : XLogRecPtr lsn)
2304 : : {
2305 : : int nchanges;
2306 : : char path[MAXPGPATH];
2307 : 31 : char *buffer = NULL;
2308 : : MemoryContext oldcxt;
2309 : : ResourceOwner oldowner;
2310 : : int fileno;
2311 : : pgoff_t offset;
2312 : :
2313 [ + + ]: 31 : if (!am_parallel_apply_worker())
2314 : 27 : maybe_start_skipping_changes(lsn);
2315 : :
2316 : : /* Make sure we have an open transaction */
2317 : 31 : begin_replication_step();
2318 : :
2319 : : /*
2320 : : * Allocate file handle and memory required to process all the messages in
2321 : : * TopTransactionContext to avoid them getting reset after each message is
2322 : : * processed.
2323 : : */
2324 : 31 : oldcxt = MemoryContextSwitchTo(TopTransactionContext);
2325 : :
2326 : : /* Open the spool file for the committed/prepared transaction */
2327 : 31 : changes_filename(path, MyLogicalRepWorker->subid, xid);
2328 [ - + ]: 31 : elog(DEBUG1, "replaying changes from file \"%s\"", path);
2329 : :
2330 : : /*
2331 : : * Make sure the file is owned by the toplevel transaction so that the
2332 : : * file will not be accidentally closed when aborting a subtransaction.
2333 : : */
2334 : 31 : oldowner = CurrentResourceOwner;
2335 : 31 : CurrentResourceOwner = TopTransactionResourceOwner;
2336 : :
2337 : 31 : stream_fd = BufFileOpenFileSet(stream_fileset, path, O_RDONLY, false);
2338 : :
2339 : 31 : CurrentResourceOwner = oldowner;
2340 : :
2341 : 31 : buffer = palloc(BLCKSZ);
2342 : :
2343 : 31 : MemoryContextSwitchTo(oldcxt);
2344 : :
2345 : : /*
2346 : : * Make sure the handle apply_dispatch methods are aware we're in a remote
2347 : : * transaction.
2348 : : */
2349 : 31 : in_remote_transaction = true;
2350 : 31 : pgstat_report_activity(STATE_RUNNING, NULL);
2351 : :
2352 : 31 : end_replication_step();
2353 : :
2354 : : /*
2355 : : * Read the entries one by one and pass them through the same logic as in
2356 : : * apply_dispatch.
2357 : : */
2358 : 31 : nchanges = 0;
2359 : : while (true)
2360 : 88470 : {
2361 : : StringInfoData s2;
2362 : : size_t nbytes;
2363 : : int len;
2364 : :
2365 [ - + ]: 88501 : CHECK_FOR_INTERRUPTS();
2366 : :
2367 : : /* read length of the on-disk record */
2368 : 88501 : nbytes = BufFileReadMaybeEOF(stream_fd, &len, sizeof(len), true);
2369 : :
2370 : : /* have we reached end of the file? */
2371 [ + + ]: 88501 : if (nbytes == 0)
2372 : 26 : break;
2373 : :
2374 : : /* do we have a correct length? */
2375 [ - + ]: 88475 : if (len <= 0)
2376 [ # # ]: 0 : elog(ERROR, "incorrect length %d in streaming transaction's changes file \"%s\"",
2377 : : len, path);
2378 : :
2379 : : /* make sure we have sufficiently large buffer */
2380 : 88475 : buffer = repalloc(buffer, len);
2381 : :
2382 : : /* and finally read the data into the buffer */
2383 : 88475 : BufFileReadExact(stream_fd, buffer, len);
2384 : :
2385 : 88475 : BufFileTell(stream_fd, &fileno, &offset);
2386 : :
2387 : : /* init a stringinfo using the buffer and call apply_dispatch */
2388 : 88475 : initReadOnlyStringInfo(&s2, buffer, len);
2389 : :
2390 : : /* Ensure we are reading the data into our memory context. */
2391 : 88475 : oldcxt = MemoryContextSwitchTo(ApplyMessageContext);
2392 : :
2393 : 88475 : apply_dispatch(&s2);
2394 : :
2395 : 88474 : MemoryContextReset(ApplyMessageContext);
2396 : :
2397 : 88474 : MemoryContextSwitchTo(oldcxt);
2398 : :
2399 : 88474 : nchanges++;
2400 : :
2401 : : /*
2402 : : * It is possible the file has been closed because we have processed
2403 : : * the transaction end message like stream_commit in which case that
2404 : : * must be the last message.
2405 : : */
2406 [ + + ]: 88474 : if (!stream_fd)
2407 : : {
2408 : 4 : ensure_last_message(stream_fileset, xid, fileno, offset);
2409 : 4 : break;
2410 : : }
2411 : :
2412 [ + + ]: 88470 : if (nchanges % 1000 == 0)
2413 [ - + ]: 83 : elog(DEBUG1, "replayed %d changes from file \"%s\"",
2414 : : nchanges, path);
2415 : : }
2416 : :
2417 [ + + ]: 30 : if (stream_fd)
2418 : 26 : stream_close_file();
2419 : :
2420 [ - + ]: 30 : elog(DEBUG1, "replayed %d (all) changes from file \"%s\"",
2421 : : nchanges, path);
2422 : :
2423 : 30 : return;
2424 : : }
2425 : :
2426 : : /*
2427 : : * Handle STREAM COMMIT message.
2428 : : */
2429 : : static void
2430 : 61 : apply_handle_stream_commit(StringInfo s)
2431 : : {
2432 : : TransactionId xid;
2433 : : LogicalRepCommitData commit_data;
2434 : : ParallelApplyWorkerInfo *winfo;
2435 : : TransApplyAction apply_action;
2436 : :
2437 : : /* Save the message before it is consumed. */
2438 : 61 : StringInfoData original_msg = *s;
2439 : :
2440 [ - + ]: 61 : if (in_streamed_transaction)
2441 [ # # ]: 0 : ereport(ERROR,
2442 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
2443 : : errmsg_internal("STREAM COMMIT message without STREAM STOP")));
2444 : :
2445 : 61 : xid = logicalrep_read_stream_commit(s, &commit_data);
2446 : 61 : set_remote_transaction_info(xid, commit_data.commit_lsn);
2447 : :
2448 : 61 : apply_action = get_transaction_apply_action(xid, &winfo);
2449 : :
2450 [ + + + + : 61 : switch (apply_action)
- ]
2451 : : {
2452 : 22 : case TRANS_LEADER_APPLY:
2453 : :
2454 : : /*
2455 : : * The transaction has been serialized to file, so replay all the
2456 : : * spooled operations.
2457 : : */
2458 : 22 : apply_spooled_messages(MyLogicalRepWorker->stream_fileset, xid,
2459 : : commit_data.commit_lsn);
2460 : :
2461 : 21 : apply_handle_commit_internal(&commit_data);
2462 : :
2463 : : /* Unlink the files with serialized changes and subxact info. */
2464 : 21 : stream_cleanup_files(MyLogicalRepWorker->subid, xid);
2465 : :
2466 [ - + ]: 21 : elog(DEBUG1, "finished processing the STREAM COMMIT command");
2467 : 21 : break;
2468 : :
2469 : 18 : case TRANS_LEADER_SEND_TO_PARALLEL:
2470 : : Assert(winfo);
2471 : :
2472 [ + - ]: 18 : if (pa_send_data(winfo, s->len, s->data))
2473 : : {
2474 : : /* Finish processing the streaming transaction. */
2475 : 18 : pa_xact_finish(winfo, commit_data.end_lsn);
2476 : 17 : break;
2477 : : }
2478 : :
2479 : : /*
2480 : : * Switch to serialize mode when we are not able to send the
2481 : : * change to parallel apply worker.
2482 : : */
2483 : 0 : pa_switch_to_partial_serialize(winfo, true);
2484 : :
2485 : : pg_fallthrough;
2486 : 2 : case TRANS_LEADER_PARTIAL_SERIALIZE:
2487 : : Assert(winfo);
2488 : :
2489 : 2 : stream_open_and_write_change(xid, LOGICAL_REP_MSG_STREAM_COMMIT,
2490 : : &original_msg);
2491 : :
2492 : 2 : pa_set_fileset_state(winfo->shared, FS_SERIALIZE_DONE);
2493 : :
2494 : : /* Finish processing the streaming transaction. */
2495 : 2 : pa_xact_finish(winfo, commit_data.end_lsn);
2496 : 2 : break;
2497 : :
2498 : 19 : case TRANS_PARALLEL_APPLY:
2499 : :
2500 : : /*
2501 : : * If the parallel apply worker is applying spooled messages then
2502 : : * close the file before committing.
2503 : : */
2504 [ + + ]: 19 : if (stream_fd)
2505 : 2 : stream_close_file();
2506 : :
2507 : 19 : apply_handle_commit_internal(&commit_data);
2508 : :
2509 : 19 : MyParallelShared->last_commit_end = XactLastCommitEnd;
2510 : :
2511 : : /*
2512 : : * It is important to set the transaction state as finished before
2513 : : * releasing the lock. See pa_wait_for_xact_finish.
2514 : : */
2515 : 19 : pa_set_xact_state(MyParallelShared, PARALLEL_TRANS_FINISHED);
2516 : 19 : pa_unlock_transaction(xid, AccessExclusiveLock);
2517 : :
2518 : 19 : pa_reset_subtrans();
2519 : :
2520 [ + + ]: 19 : elog(DEBUG1, "finished processing the STREAM COMMIT command");
2521 : 19 : break;
2522 : :
2523 : 0 : default:
2524 [ # # ]: 0 : elog(ERROR, "unexpected apply action: %d", (int) apply_action);
2525 : : break;
2526 : : }
2527 : :
2528 : : /*
2529 : : * Process any tables that are being synchronized in parallel, as well as
2530 : : * any newly added tables or sequences.
2531 : : */
2532 : 59 : ProcessSyncingRelations(commit_data.end_lsn);
2533 : :
2534 : 59 : pgstat_report_activity(STATE_IDLE, NULL);
2535 : :
2536 : 59 : reset_apply_remote_context();
2537 : 59 : }
2538 : :
2539 : : /*
2540 : : * Helper function for apply_handle_commit and apply_handle_stream_commit.
2541 : : */
2542 : : static void
2543 : 499 : apply_handle_commit_internal(LogicalRepCommitData *commit_data)
2544 : : {
2545 [ + + ]: 499 : if (is_skipping_changes())
2546 : : {
2547 : 2 : stop_skipping_changes();
2548 : :
2549 : : /*
2550 : : * Start a new transaction to clear the subskiplsn, if not started
2551 : : * yet.
2552 : : */
2553 [ + + ]: 2 : if (!IsTransactionState())
2554 : 1 : StartTransactionCommand();
2555 : : }
2556 : :
2557 [ + - ]: 499 : if (IsTransactionState())
2558 : : {
2559 : : /*
2560 : : * The transaction is either non-empty or skipped, so we clear the
2561 : : * subskiplsn.
2562 : : */
2563 : 499 : clear_subscription_skip_lsn(commit_data->commit_lsn);
2564 : :
2565 : : /*
2566 : : * Update origin state so we can restart streaming from correct
2567 : : * position in case of crash.
2568 : : */
2569 : 499 : replorigin_xact_state.origin_lsn = commit_data->end_lsn;
2570 : 499 : replorigin_xact_state.origin_timestamp = commit_data->committime;
2571 : :
2572 : 499 : CommitTransactionCommand();
2573 : :
2574 [ + + ]: 499 : if (IsTransactionBlock())
2575 : : {
2576 : 4 : EndTransactionBlock(false);
2577 : 4 : CommitTransactionCommand();
2578 : : }
2579 : :
2580 : 499 : pgstat_report_stat(false);
2581 : :
2582 : 499 : store_flush_position(commit_data->end_lsn, XactLastCommitEnd);
2583 : : }
2584 : : else
2585 : : {
2586 : : /* Process any invalidation messages that might have accumulated. */
2587 : 0 : AcceptInvalidationMessages();
2588 : 0 : maybe_reread_subscription();
2589 : : }
2590 : :
2591 : 499 : in_remote_transaction = false;
2592 : 499 : }
2593 : :
2594 : : /*
2595 : : * Handle RELATION message.
2596 : : *
2597 : : * Note we don't do validation against local schema here. The validation
2598 : : * against local schema is postponed until first change for given relation
2599 : : * comes as we only care about it when applying changes for it anyway and we
2600 : : * do less locking this way.
2601 : : */
2602 : : static void
2603 : 537 : apply_handle_relation(StringInfo s)
2604 : : {
2605 : : LogicalRepRelation *rel;
2606 : :
2607 [ + + ]: 537 : if (handle_streamed_transaction(LOGICAL_REP_MSG_RELATION, s))
2608 : 36 : return;
2609 : :
2610 : 501 : rel = logicalrep_read_rel(s);
2611 : 501 : logicalrep_relmap_update(rel);
2612 : :
2613 : : /* Also reset all entries in the partition map that refer to remoterel. */
2614 : 501 : logicalrep_partmap_reset_relmap(rel);
2615 : : }
2616 : :
2617 : : /*
2618 : : * Handle TYPE message.
2619 : : *
2620 : : * This implementation pays no attention to TYPE messages; we expect the user
2621 : : * to have set things up so that the incoming data is acceptable to the input
2622 : : * functions for the locally subscribed tables. Hence, we just read and
2623 : : * discard the message.
2624 : : */
2625 : : static void
2626 : 18 : apply_handle_type(StringInfo s)
2627 : : {
2628 : : LogicalRepTyp typ;
2629 : :
2630 [ - + ]: 18 : if (handle_streamed_transaction(LOGICAL_REP_MSG_TYPE, s))
2631 : 0 : return;
2632 : :
2633 : 18 : logicalrep_read_typ(s, &typ);
2634 : : }
2635 : :
2636 : : /*
2637 : : * Check that we (the subscription owner) have sufficient privileges on the
2638 : : * target relation to perform the given operation.
2639 : : */
2640 : : static void
2641 : 241246 : TargetPrivilegesCheck(Relation rel, AclMode mode)
2642 : : {
2643 : : Oid relid;
2644 : : AclResult aclresult;
2645 : :
2646 : 241246 : relid = RelationGetRelid(rel);
2647 : 241246 : aclresult = pg_class_aclcheck(relid, GetUserId(), mode);
2648 [ + + ]: 241246 : if (aclresult != ACLCHECK_OK)
2649 : 12 : aclcheck_error(aclresult,
2650 : 12 : get_relkind_objtype(rel->rd_rel->relkind),
2651 : 12 : get_rel_name(relid));
2652 : :
2653 : : /*
2654 : : * We lack the infrastructure to honor RLS policies. It might be possible
2655 : : * to add such infrastructure here, but tablesync workers lack it, too, so
2656 : : * we don't bother. RLS does not ordinarily apply to TRUNCATE commands,
2657 : : * but it seems dangerous to replicate a TRUNCATE and then refuse to
2658 : : * replicate subsequent INSERTs, so we forbid all commands the same.
2659 : : */
2660 [ + + ]: 241234 : if (check_enable_rls(relid, InvalidOid, false) == RLS_ENABLED)
2661 [ + - ]: 4 : ereport(ERROR,
2662 : : (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
2663 : : errmsg("user \"%s\" cannot replicate into relation with row-level security enabled: \"%s\"",
2664 : : GetUserNameFromId(GetUserId(), true),
2665 : : RelationGetRelationName(rel))));
2666 : 241230 : }
2667 : :
2668 : : /*
2669 : : * Handle INSERT message.
2670 : : */
2671 : :
2672 : : static void
2673 : 206768 : apply_handle_insert(StringInfo s)
2674 : : {
2675 : : LogicalRepRelMapEntry *rel;
2676 : : LogicalRepTupleData newtup;
2677 : : LogicalRepRelId relid;
2678 : : UserContext ucxt;
2679 : : ApplyExecutionData *edata;
2680 : : EState *estate;
2681 : : TupleTableSlot *remoteslot;
2682 : : MemoryContext oldctx;
2683 : : bool run_as_owner;
2684 : :
2685 : : /*
2686 : : * Quick return if we are skipping data modification changes or handling
2687 : : * streamed transactions.
2688 : : */
2689 [ + + + + ]: 403535 : if (is_skipping_changes() ||
2690 : 196767 : handle_streamed_transaction(LOGICAL_REP_MSG_INSERT, s))
2691 : 110057 : return;
2692 : :
2693 : 96759 : begin_replication_step();
2694 : :
2695 : 96757 : relid = logicalrep_read_insert(s, &newtup);
2696 : 96757 : rel = logicalrep_rel_open(relid, RowExclusiveLock);
2697 [ + + ]: 96745 : if (!should_apply_changes_for_rel(rel))
2698 : : {
2699 : : /*
2700 : : * The relation can't become interesting in the middle of the
2701 : : * transaction so it's safe to unlock it.
2702 : : */
2703 : 48 : logicalrep_rel_close(rel, RowExclusiveLock);
2704 : 48 : end_replication_step();
2705 : 48 : return;
2706 : : }
2707 : :
2708 : : /*
2709 : : * Make sure that any user-supplied code runs as the table owner, unless
2710 : : * the user has opted out of that behavior.
2711 : : */
2712 : 96697 : run_as_owner = MySubscription->runasowner;
2713 [ + + ]: 96697 : if (!run_as_owner)
2714 : 96687 : SwitchToUntrustedUser(rel->localrel->rd_rel->relowner, &ucxt);
2715 : :
2716 : : /* Set relation for error callback */
2717 : 96697 : remote_ctx.rel = rel;
2718 : :
2719 : : /* Initialize the executor state. */
2720 : 96697 : edata = create_edata_for_relation(rel);
2721 : 96697 : estate = edata->estate;
2722 : 96697 : remoteslot = ExecInitExtraTupleSlot(estate,
2723 : 96697 : RelationGetDescr(rel->localrel),
2724 : : &TTSOpsVirtual);
2725 : :
2726 : : /* Process and store remote tuple in the slot */
2727 [ - + ]: 96697 : oldctx = MemoryContextSwitchTo(GetPerTupleMemoryContext(estate));
2728 : 96697 : slot_store_data(remoteslot, rel, &newtup);
2729 : 96697 : slot_fill_defaults(rel, estate, remoteslot);
2730 : 96697 : MemoryContextSwitchTo(oldctx);
2731 : :
2732 : : /* For a partitioned table, insert the tuple into a partition. */
2733 [ + + ]: 96697 : if (rel->localrel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE)
2734 : 90 : apply_handle_tuple_routing(edata,
2735 : : remoteslot, NULL, CMD_INSERT);
2736 : : else
2737 : : {
2738 : 96607 : ResultRelInfo *relinfo = edata->targetRelInfo;
2739 : :
2740 : 96607 : ExecOpenIndices(relinfo, false);
2741 : 96607 : apply_handle_insert_internal(edata, relinfo, remoteslot);
2742 : 96589 : ExecCloseIndices(relinfo);
2743 : : }
2744 : :
2745 : 96633 : finish_edata(edata);
2746 : :
2747 : : /* Reset relation for error callback */
2748 : 96633 : remote_ctx.rel = NULL;
2749 : :
2750 [ + + ]: 96633 : if (!run_as_owner)
2751 : 96628 : RestoreUserContext(&ucxt);
2752 : :
2753 : 96633 : logicalrep_rel_close(rel, NoLock);
2754 : :
2755 : 96633 : end_replication_step();
2756 : : }
2757 : :
2758 : : /*
2759 : : * Workhorse for apply_handle_insert()
2760 : : * relinfo is for the relation we're actually inserting into
2761 : : * (could be a child partition of edata->targetRelInfo)
2762 : : */
2763 : : static void
2764 : 96698 : apply_handle_insert_internal(ApplyExecutionData *edata,
2765 : : ResultRelInfo *relinfo,
2766 : : TupleTableSlot *remoteslot)
2767 : : {
2768 : 96698 : EState *estate = edata->estate;
2769 : :
2770 : : /* Caller should have opened indexes already. */
2771 : : Assert(relinfo->ri_IndexRelationDescs != NULL ||
2772 : : !relinfo->ri_RelationDesc->rd_rel->relhasindex ||
2773 : : RelationGetIndexList(relinfo->ri_RelationDesc) == NIL);
2774 : :
2775 : : /* Caller will not have done this bit. */
2776 : : Assert(relinfo->ri_onConflictArbiterIndexes == NIL);
2777 : 96698 : InitConflictIndexes(relinfo);
2778 : :
2779 : : /* Do the insert. */
2780 : 96698 : TargetPrivilegesCheck(relinfo->ri_RelationDesc, ACL_INSERT);
2781 : 96689 : ExecSimpleRelationInsert(relinfo, estate, remoteslot);
2782 : 96634 : }
2783 : :
2784 : : /*
2785 : : * Check if the logical replication relation is updatable and throw
2786 : : * appropriate error if it isn't.
2787 : : */
2788 : : static void
2789 : 72306 : check_relation_updatable(LogicalRepRelMapEntry *rel)
2790 : : {
2791 : : /*
2792 : : * For partitioned tables, we only need to care if the target partition is
2793 : : * updatable (aka has PK or RI defined for it).
2794 : : */
2795 [ + + ]: 72306 : if (rel->localrel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE)
2796 : 30 : return;
2797 : :
2798 : : /* Updatable, no error. */
2799 [ + + ]: 72276 : if (rel->updatable)
2800 : 72274 : return;
2801 : :
2802 : : /* Use the entry, so this matches what updatable was decided from. */
2803 [ - + ]: 2 : if (rel->idxisreplident)
2804 : : {
2805 [ # # ]: 0 : ereport(ERROR,
2806 : : (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
2807 : : errmsg("publisher did not send replica identity column "
2808 : : "expected by the logical replication target relation \"%s.%s\"",
2809 : : rel->remoterel.nspname, rel->remoterel.relname)));
2810 : : }
2811 : :
2812 [ + - ]: 2 : ereport(ERROR,
2813 : : (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
2814 : : errmsg("logical replication target relation \"%s.%s\" has "
2815 : : "neither REPLICA IDENTITY index nor PRIMARY "
2816 : : "KEY and published relation does not have "
2817 : : "REPLICA IDENTITY FULL",
2818 : : rel->remoterel.nspname, rel->remoterel.relname)));
2819 : : }
2820 : :
2821 : : /*
2822 : : * Handle UPDATE message.
2823 : : *
2824 : : * TODO: FDW support
2825 : : */
2826 : : static void
2827 : 66179 : apply_handle_update(StringInfo s)
2828 : : {
2829 : : LogicalRepRelMapEntry *rel;
2830 : : LogicalRepRelId relid;
2831 : : UserContext ucxt;
2832 : : ApplyExecutionData *edata;
2833 : : EState *estate;
2834 : : LogicalRepTupleData oldtup;
2835 : : LogicalRepTupleData newtup;
2836 : : bool has_oldtup;
2837 : : TupleTableSlot *remoteslot;
2838 : : RTEPermissionInfo *target_perminfo;
2839 : : MemoryContext oldctx;
2840 : : bool run_as_owner;
2841 : :
2842 : : /*
2843 : : * Quick return if we are skipping data modification changes or handling
2844 : : * streamed transactions.
2845 : : */
2846 [ + + + + ]: 132355 : if (is_skipping_changes() ||
2847 : 66176 : handle_streamed_transaction(LOGICAL_REP_MSG_UPDATE, s))
2848 : 34223 : return;
2849 : :
2850 : 31956 : begin_replication_step();
2851 : :
2852 : 31955 : relid = logicalrep_read_update(s, &has_oldtup, &oldtup,
2853 : : &newtup);
2854 : 31955 : rel = logicalrep_rel_open(relid, RowExclusiveLock);
2855 [ - + ]: 31955 : if (!should_apply_changes_for_rel(rel))
2856 : : {
2857 : : /*
2858 : : * The relation can't become interesting in the middle of the
2859 : : * transaction so it's safe to unlock it.
2860 : : */
2861 : 0 : logicalrep_rel_close(rel, RowExclusiveLock);
2862 : 0 : end_replication_step();
2863 : 0 : return;
2864 : : }
2865 : :
2866 : : /* Set relation for error callback */
2867 : 31955 : remote_ctx.rel = rel;
2868 : :
2869 : : /* Check if we can do the update. */
2870 : 31955 : check_relation_updatable(rel);
2871 : :
2872 : : /*
2873 : : * Make sure that any user-supplied code runs as the table owner, unless
2874 : : * the user has opted out of that behavior.
2875 : : */
2876 : 31953 : run_as_owner = MySubscription->runasowner;
2877 [ + + ]: 31953 : if (!run_as_owner)
2878 : 31949 : SwitchToUntrustedUser(rel->localrel->rd_rel->relowner, &ucxt);
2879 : :
2880 : : /* Initialize the executor state. */
2881 : 31952 : edata = create_edata_for_relation(rel);
2882 : 31952 : estate = edata->estate;
2883 : 31952 : remoteslot = ExecInitExtraTupleSlot(estate,
2884 : 31952 : RelationGetDescr(rel->localrel),
2885 : : &TTSOpsVirtual);
2886 : :
2887 : : /*
2888 : : * Populate updatedCols so that per-column triggers can fire, and so
2889 : : * executor can correctly pass down indexUnchanged hint. This could
2890 : : * include more columns than were actually changed on the publisher
2891 : : * because the logical replication protocol doesn't contain that
2892 : : * information. But it would for example exclude columns that only exist
2893 : : * on the subscriber, since we are not touching those.
2894 : : */
2895 : 31952 : target_perminfo = list_nth(estate->es_rteperminfos, 0);
2896 [ + + ]: 159363 : for (int i = 0; i < remoteslot->tts_tupleDescriptor->natts; i++)
2897 : : {
2898 : 127411 : CompactAttribute *att = TupleDescCompactAttr(remoteslot->tts_tupleDescriptor, i);
2899 : 127411 : int remoteattnum = rel->attrmap->attnums[i];
2900 : :
2901 [ + + + + ]: 127411 : if (!att->attisdropped && remoteattnum >= 0)
2902 : : {
2903 [ - + ]: 68894 : if (remoteattnum >= newtup.ncols)
2904 [ # # ]: 0 : ereport(ERROR,
2905 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
2906 : : errmsg_plural("logical replication column %d not found in tuple: only %d column received",
2907 : : "logical replication column %d not found in tuple: only %d columns received",
2908 : : newtup.ncols,
2909 : : remoteattnum + 1, newtup.ncols)));
2910 : :
2911 [ + - ]: 68894 : if (newtup.colstatus[remoteattnum] != LOGICALREP_COLUMN_UNCHANGED)
2912 : 68894 : target_perminfo->updatedCols =
2913 : 68894 : bms_add_member(target_perminfo->updatedCols,
2914 : : i + 1 - FirstLowInvalidHeapAttributeNumber);
2915 : : }
2916 : : }
2917 : :
2918 : : /* Build the search tuple. */
2919 [ - + ]: 31952 : oldctx = MemoryContextSwitchTo(GetPerTupleMemoryContext(estate));
2920 : 31952 : slot_store_data(remoteslot, rel,
2921 [ + + ]: 31952 : has_oldtup ? &oldtup : &newtup);
2922 : 31952 : MemoryContextSwitchTo(oldctx);
2923 : :
2924 : : /* For a partitioned table, apply update to correct partition. */
2925 [ + + ]: 31952 : if (rel->localrel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE)
2926 : 13 : apply_handle_tuple_routing(edata,
2927 : : remoteslot, &newtup, CMD_UPDATE);
2928 : : else
2929 : 31939 : apply_handle_update_internal(edata, edata->targetRelInfo,
2930 : : remoteslot, &newtup);
2931 : :
2932 : 31943 : finish_edata(edata);
2933 : :
2934 : : /* Reset relation for error callback */
2935 : 31943 : remote_ctx.rel = NULL;
2936 : :
2937 [ + + ]: 31943 : if (!run_as_owner)
2938 : 31941 : RestoreUserContext(&ucxt);
2939 : :
2940 : 31943 : logicalrep_rel_close(rel, NoLock);
2941 : :
2942 : 31943 : end_replication_step();
2943 : : }
2944 : :
2945 : : /*
2946 : : * Workhorse for apply_handle_update()
2947 : : * relinfo is for the relation we're actually updating in
2948 : : * (could be a child partition of edata->targetRelInfo)
2949 : : */
2950 : : static void
2951 : 31939 : apply_handle_update_internal(ApplyExecutionData *edata,
2952 : : ResultRelInfo *relinfo,
2953 : : TupleTableSlot *remoteslot,
2954 : : LogicalRepTupleData *newtup)
2955 : : {
2956 : 31939 : EState *estate = edata->estate;
2957 : 31939 : LogicalRepRelMapEntry *relmapentry = edata->targetRel;
2958 : 31939 : Relation localrel = relinfo->ri_RelationDesc;
2959 : : EPQState epqstate;
2960 : 31939 : TupleTableSlot *localslot = NULL;
2961 : 31939 : ConflictTupleInfo conflicttuple = {0};
2962 : : bool found;
2963 : : MemoryContext oldctx;
2964 : :
2965 : 31939 : EvalPlanQualInit(&epqstate, estate, NULL, NIL, -1, NIL);
2966 : :
2967 : 31939 : INJECTION_POINT("apply-update-before-open-indices", NULL);
2968 : :
2969 : 31939 : ExecOpenIndices(relinfo, false);
2970 : :
2971 : 31939 : found = FindReplTupleInLocalRel(edata, localrel, relmapentry,
2972 : : remoteslot, &localslot);
2973 : :
2974 : : /*
2975 : : * Tuple found.
2976 : : *
2977 : : * Note this will fail if there are other conflicting unique indexes.
2978 : : */
2979 [ + + ]: 31932 : if (found)
2980 : : {
2981 : : /*
2982 : : * Report the conflict if the tuple was modified by a different
2983 : : * origin.
2984 : : */
2985 [ + + ]: 31916 : if (GetTupleTransactionInfo(localslot, &conflicttuple.xmin,
2986 : 2 : &conflicttuple.origin, &conflicttuple.ts) &&
2987 [ + - ]: 2 : conflicttuple.origin != replorigin_xact_state.origin)
2988 : : {
2989 : : TupleTableSlot *newslot;
2990 : :
2991 : : /* Store the new tuple for conflict reporting */
2992 : 2 : newslot = table_slot_create(localrel, &estate->es_tupleTable);
2993 : 2 : slot_store_data(newslot, relmapentry, newtup);
2994 : :
2995 : 2 : conflicttuple.slot = localslot;
2996 : :
2997 : 2 : ReportApplyConflict(estate, relinfo, LOG, CT_UPDATE_ORIGIN_DIFFERS,
2998 : : remoteslot, newslot,
2999 : : list_make1(&conflicttuple));
3000 : : }
3001 : :
3002 : : /* Process and store remote tuple in the slot */
3003 [ + - ]: 31916 : oldctx = MemoryContextSwitchTo(GetPerTupleMemoryContext(estate));
3004 : 31916 : slot_modify_data(remoteslot, localslot, relmapentry, newtup);
3005 : 31916 : MemoryContextSwitchTo(oldctx);
3006 : :
3007 : 31916 : EvalPlanQualSetSlot(&epqstate, remoteslot);
3008 : :
3009 : 31916 : InitConflictIndexes(relinfo);
3010 : :
3011 : : /* Do the actual update. */
3012 : 31916 : TargetPrivilegesCheck(relinfo->ri_RelationDesc, ACL_UPDATE);
3013 : 31916 : ExecSimpleRelationUpdate(relinfo, estate, &epqstate, localslot,
3014 : : remoteslot);
3015 : : }
3016 : : else
3017 : : {
3018 : : ConflictType type;
3019 : 16 : TupleTableSlot *newslot = localslot;
3020 : :
3021 : : /*
3022 : : * Detecting whether the tuple was recently deleted or never existed
3023 : : * is crucial to avoid misleading the user during conflict handling.
3024 : : */
3025 [ + + ]: 16 : if (FindDeletedTupleInLocalRel(localrel, relmapentry, remoteslot,
3026 : : &conflicttuple.xmin,
3027 : : &conflicttuple.origin,
3028 : 3 : &conflicttuple.ts) &&
3029 [ + - ]: 3 : conflicttuple.origin != replorigin_xact_state.origin)
3030 : 3 : type = CT_UPDATE_DELETED;
3031 : : else
3032 : 13 : type = CT_UPDATE_MISSING;
3033 : :
3034 : : /* Store the new tuple for conflict reporting */
3035 : 16 : slot_store_data(newslot, relmapentry, newtup);
3036 : :
3037 : : /*
3038 : : * The tuple to be updated could not be found or was deleted. Do
3039 : : * nothing except for emitting a log message.
3040 : : */
3041 : 16 : ReportApplyConflict(estate, relinfo, LOG, type, remoteslot, newslot,
3042 : : list_make1(&conflicttuple));
3043 : : }
3044 : :
3045 : : /* Cleanup. */
3046 : 31930 : ExecCloseIndices(relinfo);
3047 : 31930 : EvalPlanQualEnd(&epqstate);
3048 : 31930 : }
3049 : :
3050 : : /*
3051 : : * Handle DELETE message.
3052 : : *
3053 : : * TODO: FDW support
3054 : : */
3055 : : static void
3056 : 81936 : apply_handle_delete(StringInfo s)
3057 : : {
3058 : : LogicalRepRelMapEntry *rel;
3059 : : LogicalRepTupleData oldtup;
3060 : : LogicalRepRelId relid;
3061 : : UserContext ucxt;
3062 : : ApplyExecutionData *edata;
3063 : : EState *estate;
3064 : : TupleTableSlot *remoteslot;
3065 : : MemoryContext oldctx;
3066 : : bool run_as_owner;
3067 : :
3068 : : /*
3069 : : * Quick return if we are skipping data modification changes or handling
3070 : : * streamed transactions.
3071 : : */
3072 [ + - + + ]: 163872 : if (is_skipping_changes() ||
3073 : 81936 : handle_streamed_transaction(LOGICAL_REP_MSG_DELETE, s))
3074 : 41615 : return;
3075 : :
3076 : 40321 : begin_replication_step();
3077 : :
3078 : 40321 : relid = logicalrep_read_delete(s, &oldtup);
3079 : 40321 : rel = logicalrep_rel_open(relid, RowExclusiveLock);
3080 [ - + ]: 40321 : if (!should_apply_changes_for_rel(rel))
3081 : : {
3082 : : /*
3083 : : * The relation can't become interesting in the middle of the
3084 : : * transaction so it's safe to unlock it.
3085 : : */
3086 : 0 : logicalrep_rel_close(rel, RowExclusiveLock);
3087 : 0 : end_replication_step();
3088 : 0 : return;
3089 : : }
3090 : :
3091 : : /* Set relation for error callback */
3092 : 40321 : remote_ctx.rel = rel;
3093 : :
3094 : : /* Check if we can do the delete. */
3095 : 40321 : check_relation_updatable(rel);
3096 : :
3097 : : /*
3098 : : * Make sure that any user-supplied code runs as the table owner, unless
3099 : : * the user has opted out of that behavior.
3100 : : */
3101 : 40321 : run_as_owner = MySubscription->runasowner;
3102 [ + + ]: 40321 : if (!run_as_owner)
3103 : 40319 : SwitchToUntrustedUser(rel->localrel->rd_rel->relowner, &ucxt);
3104 : :
3105 : : /* Initialize the executor state. */
3106 : 40321 : edata = create_edata_for_relation(rel);
3107 : 40321 : estate = edata->estate;
3108 : 40321 : remoteslot = ExecInitExtraTupleSlot(estate,
3109 : 40321 : RelationGetDescr(rel->localrel),
3110 : : &TTSOpsVirtual);
3111 : :
3112 : : /* Build the search tuple. */
3113 [ - + ]: 40321 : oldctx = MemoryContextSwitchTo(GetPerTupleMemoryContext(estate));
3114 : 40321 : slot_store_data(remoteslot, rel, &oldtup);
3115 : 40321 : MemoryContextSwitchTo(oldctx);
3116 : :
3117 : : /* For a partitioned table, apply delete to correct partition. */
3118 [ + + ]: 40321 : if (rel->localrel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE)
3119 : 17 : apply_handle_tuple_routing(edata,
3120 : : remoteslot, NULL, CMD_DELETE);
3121 : : else
3122 : : {
3123 : 40304 : ResultRelInfo *relinfo = edata->targetRelInfo;
3124 : :
3125 : 40304 : ExecOpenIndices(relinfo, false);
3126 : 40304 : apply_handle_delete_internal(edata, relinfo,
3127 : : remoteslot, rel);
3128 : 40304 : ExecCloseIndices(relinfo);
3129 : : }
3130 : :
3131 : 40321 : finish_edata(edata);
3132 : :
3133 : : /* Reset relation for error callback */
3134 : 40321 : remote_ctx.rel = NULL;
3135 : :
3136 [ + + ]: 40321 : if (!run_as_owner)
3137 : 40319 : RestoreUserContext(&ucxt);
3138 : :
3139 : 40321 : logicalrep_rel_close(rel, NoLock);
3140 : :
3141 : 40321 : end_replication_step();
3142 : : }
3143 : :
3144 : : /*
3145 : : * Workhorse for apply_handle_delete()
3146 : : * relinfo is for the relation we're actually deleting from
3147 : : * (could be a child partition of edata->targetRelInfo)
3148 : : */
3149 : : static void
3150 : 40321 : apply_handle_delete_internal(ApplyExecutionData *edata,
3151 : : ResultRelInfo *relinfo,
3152 : : TupleTableSlot *remoteslot,
3153 : : LogicalRepRelMapEntry *relmapentry)
3154 : : {
3155 : 40321 : EState *estate = edata->estate;
3156 : 40321 : Relation localrel = relinfo->ri_RelationDesc;
3157 : : EPQState epqstate;
3158 : : TupleTableSlot *localslot;
3159 : 40321 : ConflictTupleInfo conflicttuple = {0};
3160 : : bool found;
3161 : :
3162 : 40321 : EvalPlanQualInit(&epqstate, estate, NULL, NIL, -1, NIL);
3163 : :
3164 : : /* Caller should have opened indexes already. */
3165 : : Assert(relinfo->ri_IndexRelationDescs != NULL ||
3166 : : !localrel->rd_rel->relhasindex ||
3167 : : RelationGetIndexList(localrel) == NIL);
3168 : :
3169 : 40321 : found = FindReplTupleInLocalRel(edata, localrel, relmapentry,
3170 : : remoteslot, &localslot);
3171 : :
3172 : : /* If found delete it. */
3173 [ + + ]: 40321 : if (found)
3174 : : {
3175 : : /*
3176 : : * Report the conflict if the tuple was modified by a different
3177 : : * origin.
3178 : : */
3179 [ + + ]: 40312 : if (GetTupleTransactionInfo(localslot, &conflicttuple.xmin,
3180 : 6 : &conflicttuple.origin, &conflicttuple.ts) &&
3181 [ + + ]: 6 : conflicttuple.origin != replorigin_xact_state.origin)
3182 : : {
3183 : 5 : conflicttuple.slot = localslot;
3184 : 5 : ReportApplyConflict(estate, relinfo, LOG, CT_DELETE_ORIGIN_DIFFERS,
3185 : : remoteslot, NULL,
3186 : : list_make1(&conflicttuple));
3187 : : }
3188 : :
3189 : 40312 : EvalPlanQualSetSlot(&epqstate, localslot);
3190 : :
3191 : : /* Do the actual delete. */
3192 : 40312 : TargetPrivilegesCheck(relinfo->ri_RelationDesc, ACL_DELETE);
3193 : 40312 : ExecSimpleRelationDelete(relinfo, estate, &epqstate, localslot);
3194 : : }
3195 : : else
3196 : : {
3197 : : /*
3198 : : * The tuple to be deleted could not be found. Do nothing except for
3199 : : * emitting a log message.
3200 : : */
3201 : 9 : ReportApplyConflict(estate, relinfo, LOG, CT_DELETE_MISSING,
3202 : : remoteslot, NULL, list_make1(&conflicttuple));
3203 : : }
3204 : :
3205 : : /* Cleanup. */
3206 : 40321 : EvalPlanQualEnd(&epqstate);
3207 : 40321 : }
3208 : :
3209 : : /*
3210 : : * Try to find a tuple received from the publication side (in 'remoteslot') in
3211 : : * the corresponding local relation using either replica identity index,
3212 : : * primary key, index or if needed, sequential scan.
3213 : : *
3214 : : * 'relmapentry' is the relation map entry for 'localrel'. It tells which
3215 : : * index to use, if any, and whether that index is the relation's replica
3216 : : * identity or primary key.
3217 : : *
3218 : : * Local tuple, if found, is returned in '*localslot'.
3219 : : */
3220 : : static bool
3221 : 72273 : FindReplTupleInLocalRel(ApplyExecutionData *edata, Relation localrel,
3222 : : LogicalRepRelMapEntry *relmapentry,
3223 : : TupleTableSlot *remoteslot,
3224 : : TupleTableSlot **localslot)
3225 : : {
3226 : 72273 : EState *estate = edata->estate;
3227 : 72273 : Oid localidxoid = relmapentry->localindexoid;
3228 : : bool found;
3229 : :
3230 : : /*
3231 : : * Regardless of the top-level operation, we're performing a read here, so
3232 : : * check for SELECT privileges.
3233 : : */
3234 : 72273 : TargetPrivilegesCheck(localrel, ACL_SELECT);
3235 : :
3236 : 72266 : *localslot = table_slot_create(localrel, &estate->es_tupleTable);
3237 : :
3238 : : Assert(OidIsValid(localidxoid) ||
3239 : : (relmapentry->remoterel.replident == REPLICA_IDENTITY_FULL));
3240 : :
3241 [ + + ]: 72266 : if (OidIsValid(localidxoid))
3242 : : {
3243 : : #ifdef USE_ASSERT_CHECKING
3244 : : Relation idxrel = index_open(localidxoid, AccessShareLock);
3245 : :
3246 : : if (relmapentry->idxisreplident)
3247 : : {
3248 : : /*
3249 : : * We cannot assert this is still the replica identity or primary
3250 : : * key. DROP INDEX CONCURRENTLY and REINDEX CONCURRENTLY clear
3251 : : * indisvalid and indisreplident without conflicting with our
3252 : : * RowExclusiveLock, so GetRelationIdentityOrPK() may no longer
3253 : : * return it. Unique and non-partial is what the scan actually
3254 : : * relies on, and no DDL can take those away.
3255 : : */
3256 : : Assert(idxrel->rd_index->indisunique);
3257 : : Assert(heap_attisnull(idxrel->rd_indextuple,
3258 : : Anum_pg_index_indpred, NULL));
3259 : : }
3260 : : else
3261 : : {
3262 : : /* Otherwise every match is compared, so we need a whole row. */
3263 : : Assert(relmapentry->remoterel.replident == REPLICA_IDENTITY_FULL);
3264 : : Assert(IsIndexUsableForReplicaIdentityFull(idxrel,
3265 : : relmapentry->attrmap));
3266 : : }
3267 : : index_close(idxrel, AccessShareLock);
3268 : : #endif
3269 : :
3270 : 72114 : found = RelationFindReplTupleByIndex(localrel, localidxoid,
3271 : 72114 : relmapentry->idxisreplident,
3272 : : LockTupleExclusive,
3273 : : remoteslot, *localslot);
3274 : : }
3275 : : else
3276 : 152 : found = RelationFindReplTupleSeq(localrel, LockTupleExclusive,
3277 : : remoteslot, *localslot);
3278 : :
3279 : 72266 : return found;
3280 : : }
3281 : :
3282 : : /*
3283 : : * Determine whether the index can reliably locate the deleted tuple in the
3284 : : * local relation.
3285 : : *
3286 : : * An index may exclude deleted tuples if it was re-indexed or re-created during
3287 : : * change application. Therefore, an index is considered usable only if the
3288 : : * conflict detection slot.xmin (conflict_detection_xmin) is greater than the
3289 : : * index tuple's xmin. This ensures that any tuples deleted prior to the index
3290 : : * creation or re-indexing are not relevant for conflict detection in the
3291 : : * current apply worker.
3292 : : *
3293 : : * Note that indexes may also be excluded if they were modified by other DDL
3294 : : * operations, such as ALTER INDEX. However, this is acceptable, as the
3295 : : * likelihood of such DDL changes coinciding with the need to scan dead
3296 : : * tuples for the update_deleted is low.
3297 : : */
3298 : : static bool
3299 : 2 : IsIndexUsableForFindingDeletedTuple(Oid localindexoid,
3300 : : TransactionId conflict_detection_xmin)
3301 : : {
3302 : : HeapTuple index_tuple;
3303 : : TransactionId index_xmin;
3304 : :
3305 : 2 : index_tuple = SearchSysCache1(INDEXRELID, ObjectIdGetDatum(localindexoid));
3306 : :
3307 [ - + ]: 2 : if (!HeapTupleIsValid(index_tuple)) /* should not happen */
3308 [ # # ]: 0 : elog(ERROR, "cache lookup failed for index %u", localindexoid);
3309 : :
3310 : : /*
3311 : : * No need to check for a frozen transaction ID, as
3312 : : * TransactionIdPrecedes() manages it internally, treating it as falling
3313 : : * behind the conflict_detection_xmin.
3314 : : */
3315 : 2 : index_xmin = HeapTupleHeaderGetXmin(index_tuple->t_data);
3316 : :
3317 : 2 : ReleaseSysCache(index_tuple);
3318 : :
3319 : 2 : return TransactionIdPrecedes(index_xmin, conflict_detection_xmin);
3320 : : }
3321 : :
3322 : : /*
3323 : : * Attempts to locate a deleted tuple in the local relation that matches the
3324 : : * values of the tuple received from the publication side (in 'remoteslot').
3325 : : * The search is performed using either the replica identity index, primary
3326 : : * key, other available index, or a sequential scan if necessary.
3327 : : *
3328 : : * 'relmapentry' is the relation map entry for 'localrel', as in
3329 : : * FindReplTupleInLocalRel().
3330 : : *
3331 : : * Returns true if the deleted tuple is found. If found, the transaction ID,
3332 : : * origin, and commit timestamp of the deletion are stored in '*delete_xid',
3333 : : * '*delete_origin', and '*delete_time' respectively.
3334 : : */
3335 : : static bool
3336 : 18 : FindDeletedTupleInLocalRel(Relation localrel,
3337 : : LogicalRepRelMapEntry *relmapentry,
3338 : : TupleTableSlot *remoteslot,
3339 : : TransactionId *delete_xid, ReplOriginId *delete_origin,
3340 : : TimestampTz *delete_time)
3341 : : {
3342 : 18 : Oid localidxoid = relmapentry->localindexoid;
3343 : : TransactionId oldestxmin;
3344 : :
3345 : : /*
3346 : : * Return false if either dead tuples are not retained or commit timestamp
3347 : : * data is not available.
3348 : : */
3349 [ + + - + ]: 18 : if (!MySubscription->retaindeadtuples || !track_commit_timestamp)
3350 : 14 : return false;
3351 : :
3352 : : /*
3353 : : * For conflict detection, we use the leader worker's
3354 : : * oldest_nonremovable_xid value instead of invoking
3355 : : * GetOldestNonRemovableTransactionId() or using the conflict detection
3356 : : * slot's xmin. The oldest_nonremovable_xid acts as a threshold to
3357 : : * identify tuples that were recently deleted. These deleted tuples are no
3358 : : * longer visible to concurrent transactions. However, if a remote update
3359 : : * matches such a tuple, we log an update_deleted conflict.
3360 : : *
3361 : : * While GetOldestNonRemovableTransactionId() and slot.xmin may return
3362 : : * transaction IDs older than oldest_nonremovable_xid, for our current
3363 : : * purpose, it is acceptable to treat tuples deleted by transactions prior
3364 : : * to oldest_nonremovable_xid as update_missing conflicts.
3365 : : */
3366 [ + - ]: 4 : if (am_leader_apply_worker())
3367 : : {
3368 : 4 : oldestxmin = MyLogicalRepWorker->oldest_nonremovable_xid;
3369 : : }
3370 : : else
3371 : : {
3372 : : LogicalRepWorker *leader;
3373 : :
3374 : : /*
3375 : : * Obtain the information from the leader apply worker as only the
3376 : : * leader manages oldest_nonremovable_xid (see
3377 : : * maybe_advance_nonremovable_xid() for details).
3378 : : */
3379 : 0 : LWLockAcquire(LogicalRepWorkerLock, LW_SHARED);
3380 : 0 : leader = logicalrep_worker_find(WORKERTYPE_APPLY,
3381 : 0 : MyLogicalRepWorker->subid, InvalidOid,
3382 : : false);
3383 [ # # ]: 0 : if (!leader)
3384 : : {
3385 [ # # ]: 0 : ereport(ERROR,
3386 : : (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
3387 : : errmsg("could not detect conflict as the leader apply worker has exited")));
3388 : : }
3389 : :
3390 : 0 : SpinLockAcquire(&leader->relmutex);
3391 : 0 : oldestxmin = leader->oldest_nonremovable_xid;
3392 : 0 : SpinLockRelease(&leader->relmutex);
3393 : 0 : LWLockRelease(LogicalRepWorkerLock);
3394 : : }
3395 : :
3396 : : /*
3397 : : * Return false if the leader apply worker has stopped retaining
3398 : : * information for detecting conflicts. This implies that update_deleted
3399 : : * can no longer be reliably detected.
3400 : : */
3401 [ - + ]: 4 : if (!TransactionIdIsValid(oldestxmin))
3402 : 0 : return false;
3403 : :
3404 [ + + + + ]: 6 : if (OidIsValid(localidxoid) &&
3405 : 2 : IsIndexUsableForFindingDeletedTuple(localidxoid, oldestxmin))
3406 : 1 : return RelationFindDeletedTupleInfoByIndex(localrel, localidxoid,
3407 : 1 : relmapentry->idxisreplident,
3408 : : remoteslot, oldestxmin,
3409 : : delete_xid, delete_origin,
3410 : : delete_time);
3411 : : else
3412 : 3 : return RelationFindDeletedTupleInfoSeq(localrel,
3413 [ - + ]: 3 : relmapentry->idxisreplident ?
3414 : : localidxoid : InvalidOid,
3415 : : remoteslot, oldestxmin,
3416 : : delete_xid, delete_origin,
3417 : : delete_time);
3418 : : }
3419 : :
3420 : : /*
3421 : : * This handles insert, update, delete on a partitioned table.
3422 : : */
3423 : : static void
3424 : 120 : apply_handle_tuple_routing(ApplyExecutionData *edata,
3425 : : TupleTableSlot *remoteslot,
3426 : : LogicalRepTupleData *newtup,
3427 : : CmdType operation)
3428 : : {
3429 : 120 : EState *estate = edata->estate;
3430 : 120 : LogicalRepRelMapEntry *relmapentry = edata->targetRel;
3431 : 120 : ResultRelInfo *relinfo = edata->targetRelInfo;
3432 : 120 : Relation parentrel = relinfo->ri_RelationDesc;
3433 : : ModifyTableState *mtstate;
3434 : : PartitionTupleRouting *proute;
3435 : : ResultRelInfo *partrelinfo;
3436 : : Relation partrel;
3437 : : TupleTableSlot *remoteslot_part;
3438 : : TupleConversionMap *map;
3439 : : MemoryContext oldctx;
3440 : 120 : LogicalRepRelMapEntry *part_entry = NULL;
3441 : 120 : AttrMap *attrmap = NULL;
3442 : :
3443 : : /* ModifyTableState is needed for ExecFindPartition(). */
3444 : 120 : edata->mtstate = mtstate = makeNode(ModifyTableState);
3445 : 120 : mtstate->ps.plan = NULL;
3446 : 120 : mtstate->ps.state = estate;
3447 : 120 : mtstate->operation = operation;
3448 : 120 : mtstate->resultRelInfo = relinfo;
3449 : :
3450 : : /* ... as is PartitionTupleRouting. */
3451 : 120 : edata->proute = proute = ExecSetupPartitionTupleRouting(estate, parentrel);
3452 : :
3453 : : /*
3454 : : * Find the partition to which the "search tuple" belongs.
3455 : : */
3456 : : Assert(remoteslot != NULL);
3457 [ + - ]: 120 : oldctx = MemoryContextSwitchTo(GetPerTupleMemoryContext(estate));
3458 : 120 : partrelinfo = ExecFindPartition(mtstate, relinfo, proute,
3459 : : remoteslot, estate);
3460 : : Assert(partrelinfo != NULL);
3461 : 120 : partrel = partrelinfo->ri_RelationDesc;
3462 : :
3463 : : /*
3464 : : * Check for supported relkind. We need this since partitions might be of
3465 : : * unsupported relkinds; and the set of partitions can change, so checking
3466 : : * at CREATE/ALTER SUBSCRIPTION would be insufficient.
3467 : : */
3468 : 120 : CheckSubscriptionRelkind(partrel->rd_rel->relkind,
3469 : 120 : relmapentry->remoterel.relkind,
3470 : 120 : get_namespace_name(RelationGetNamespace(partrel)),
3471 : 120 : RelationGetRelationName(partrel));
3472 : :
3473 : : /*
3474 : : * To perform any of the operations below, the tuple must match the
3475 : : * partition's rowtype. Convert if needed or just copy, using a dedicated
3476 : : * slot to store the tuple in any case.
3477 : : */
3478 : 120 : remoteslot_part = partrelinfo->ri_PartitionTupleSlot;
3479 [ + + ]: 120 : if (remoteslot_part == NULL)
3480 : 87 : remoteslot_part = table_slot_create(partrel, &estate->es_tupleTable);
3481 : 120 : map = ExecGetRootToChildMap(partrelinfo, estate);
3482 [ + + ]: 120 : if (map != NULL)
3483 : : {
3484 : 33 : attrmap = map->attrMap;
3485 : 33 : remoteslot_part = execute_attr_map_slot(attrmap, remoteslot,
3486 : : remoteslot_part);
3487 : : }
3488 : : else
3489 : : {
3490 : 87 : remoteslot_part = ExecCopySlot(remoteslot_part, remoteslot);
3491 : 87 : slot_getallattrs(remoteslot_part);
3492 : : }
3493 : 120 : MemoryContextSwitchTo(oldctx);
3494 : :
3495 : : /* Check if we can do the update or delete on the leaf partition. */
3496 [ + + + + ]: 120 : if (operation == CMD_UPDATE || operation == CMD_DELETE)
3497 : : {
3498 : 30 : part_entry = logicalrep_partition_open(relmapentry, partrel,
3499 : : attrmap);
3500 : 30 : check_relation_updatable(part_entry);
3501 : : }
3502 : :
3503 [ + + + - ]: 120 : switch (operation)
3504 : : {
3505 : 90 : case CMD_INSERT:
3506 : 90 : apply_handle_insert_internal(edata, partrelinfo,
3507 : : remoteslot_part);
3508 : 44 : break;
3509 : :
3510 : 17 : case CMD_DELETE:
3511 : 17 : apply_handle_delete_internal(edata, partrelinfo,
3512 : : remoteslot_part, part_entry);
3513 : 17 : break;
3514 : :
3515 : 13 : case CMD_UPDATE:
3516 : :
3517 : : /*
3518 : : * For UPDATE, depending on whether or not the updated tuple
3519 : : * satisfies the partition's constraint, perform a simple UPDATE
3520 : : * of the partition or move the updated tuple into a different
3521 : : * suitable partition.
3522 : : */
3523 : : {
3524 : : TupleTableSlot *localslot;
3525 : : ResultRelInfo *partrelinfo_new;
3526 : : Relation partrel_new;
3527 : : bool found;
3528 : : EPQState epqstate;
3529 : 13 : ConflictTupleInfo conflicttuple = {0};
3530 : :
3531 : : /* Get the matching local tuple from the partition. */
3532 : 13 : found = FindReplTupleInLocalRel(edata, partrel, part_entry,
3533 : : remoteslot_part, &localslot);
3534 [ + + ]: 13 : if (!found)
3535 : : {
3536 : : ConflictType type;
3537 : 2 : TupleTableSlot *newslot = localslot;
3538 : :
3539 : : /*
3540 : : * Detecting whether the tuple was recently deleted or
3541 : : * never existed is crucial to avoid misleading the user
3542 : : * during conflict handling.
3543 : : */
3544 [ - + ]: 2 : if (FindDeletedTupleInLocalRel(partrel, part_entry,
3545 : : remoteslot_part,
3546 : : &conflicttuple.xmin,
3547 : : &conflicttuple.origin,
3548 : 0 : &conflicttuple.ts) &&
3549 [ # # ]: 0 : conflicttuple.origin != replorigin_xact_state.origin)
3550 : 0 : type = CT_UPDATE_DELETED;
3551 : : else
3552 : 2 : type = CT_UPDATE_MISSING;
3553 : :
3554 : : /* Store the new tuple for conflict reporting */
3555 : 2 : slot_store_data(newslot, part_entry, newtup);
3556 : :
3557 : : /*
3558 : : * The tuple to be updated could not be found or was
3559 : : * deleted. Do nothing except for emitting a log message.
3560 : : */
3561 : 2 : ReportApplyConflict(estate, partrelinfo, LOG,
3562 : : type, remoteslot_part, newslot,
3563 : : list_make1(&conflicttuple));
3564 : :
3565 : 2 : return;
3566 : : }
3567 : :
3568 : : /*
3569 : : * Report the conflict if the tuple was modified by a
3570 : : * different origin.
3571 : : */
3572 [ + + ]: 11 : if (GetTupleTransactionInfo(localslot, &conflicttuple.xmin,
3573 : : &conflicttuple.origin,
3574 : 1 : &conflicttuple.ts) &&
3575 [ + - ]: 1 : conflicttuple.origin != replorigin_xact_state.origin)
3576 : : {
3577 : : TupleTableSlot *newslot;
3578 : :
3579 : : /* Store the new tuple for conflict reporting */
3580 : 1 : newslot = table_slot_create(partrel, &estate->es_tupleTable);
3581 : 1 : slot_store_data(newslot, part_entry, newtup);
3582 : :
3583 : 1 : conflicttuple.slot = localslot;
3584 : :
3585 : 1 : ReportApplyConflict(estate, partrelinfo, LOG, CT_UPDATE_ORIGIN_DIFFERS,
3586 : : remoteslot_part, newslot,
3587 : : list_make1(&conflicttuple));
3588 : : }
3589 : :
3590 : : /*
3591 : : * Apply the update to the local tuple, putting the result in
3592 : : * remoteslot_part.
3593 : : */
3594 [ + - ]: 11 : oldctx = MemoryContextSwitchTo(GetPerTupleMemoryContext(estate));
3595 : 11 : slot_modify_data(remoteslot_part, localslot, part_entry,
3596 : : newtup);
3597 : 11 : MemoryContextSwitchTo(oldctx);
3598 : :
3599 : 11 : EvalPlanQualInit(&epqstate, estate, NULL, NIL, -1, NIL);
3600 : :
3601 : : /*
3602 : : * Does the updated tuple still satisfy the current
3603 : : * partition's constraint?
3604 : : */
3605 [ + - + + ]: 22 : if (!partrel->rd_rel->relispartition ||
3606 : 11 : ExecPartitionCheck(partrelinfo, remoteslot_part, estate,
3607 : : false))
3608 : : {
3609 : : /*
3610 : : * Yes, so simply UPDATE the partition. We don't call
3611 : : * apply_handle_update_internal() here, which would
3612 : : * normally do the following work, to avoid repeating some
3613 : : * work already done above to find the local tuple in the
3614 : : * partition.
3615 : : */
3616 : 10 : InitConflictIndexes(partrelinfo);
3617 : :
3618 : 10 : EvalPlanQualSetSlot(&epqstate, remoteslot_part);
3619 : 10 : TargetPrivilegesCheck(partrelinfo->ri_RelationDesc,
3620 : : ACL_UPDATE);
3621 : 10 : ExecSimpleRelationUpdate(partrelinfo, estate, &epqstate,
3622 : : localslot, remoteslot_part);
3623 : : }
3624 : : else
3625 : : {
3626 : : /* Move the tuple into the new partition. */
3627 : :
3628 : : /*
3629 : : * New partition will be found using tuple routing, which
3630 : : * can only occur via the parent table. We might need to
3631 : : * convert the tuple to the parent's rowtype. Note that
3632 : : * this is the tuple found in the partition, not the
3633 : : * original search tuple received by this function.
3634 : : */
3635 [ + - ]: 1 : if (map)
3636 : : {
3637 : : TupleConversionMap *PartitionToRootMap =
3638 : 1 : convert_tuples_by_name(RelationGetDescr(partrel),
3639 : : RelationGetDescr(parentrel));
3640 : :
3641 : : remoteslot =
3642 : 1 : execute_attr_map_slot(PartitionToRootMap->attrMap,
3643 : : remoteslot_part, remoteslot);
3644 : : }
3645 : : else
3646 : : {
3647 : 0 : remoteslot = ExecCopySlot(remoteslot, remoteslot_part);
3648 : 0 : slot_getallattrs(remoteslot);
3649 : : }
3650 : :
3651 : : /* Find the new partition. */
3652 [ + - ]: 1 : oldctx = MemoryContextSwitchTo(GetPerTupleMemoryContext(estate));
3653 : 1 : partrelinfo_new = ExecFindPartition(mtstate, relinfo,
3654 : : proute, remoteslot,
3655 : : estate);
3656 : 1 : MemoryContextSwitchTo(oldctx);
3657 : : Assert(partrelinfo_new != partrelinfo);
3658 : 1 : partrel_new = partrelinfo_new->ri_RelationDesc;
3659 : :
3660 : : /* Check that new partition also has supported relkind. */
3661 : 1 : CheckSubscriptionRelkind(partrel_new->rd_rel->relkind,
3662 : 1 : relmapentry->remoterel.relkind,
3663 : 1 : get_namespace_name(RelationGetNamespace(partrel_new)),
3664 : 1 : RelationGetRelationName(partrel_new));
3665 : :
3666 : : /* DELETE old tuple found in the old partition. */
3667 : 1 : EvalPlanQualSetSlot(&epqstate, localslot);
3668 : 1 : TargetPrivilegesCheck(partrelinfo->ri_RelationDesc, ACL_DELETE);
3669 : 1 : ExecSimpleRelationDelete(partrelinfo, estate, &epqstate, localslot);
3670 : :
3671 : : /* INSERT new tuple into the new partition. */
3672 : :
3673 : : /*
3674 : : * Convert the replacement tuple to match the destination
3675 : : * partition rowtype.
3676 : : */
3677 [ + - ]: 1 : oldctx = MemoryContextSwitchTo(GetPerTupleMemoryContext(estate));
3678 : 1 : remoteslot_part = partrelinfo_new->ri_PartitionTupleSlot;
3679 [ + - ]: 1 : if (remoteslot_part == NULL)
3680 : 1 : remoteslot_part = table_slot_create(partrel_new,
3681 : : &estate->es_tupleTable);
3682 : 1 : map = ExecGetRootToChildMap(partrelinfo_new, estate);
3683 [ - + ]: 1 : if (map != NULL)
3684 : : {
3685 : 0 : remoteslot_part = execute_attr_map_slot(map->attrMap,
3686 : : remoteslot,
3687 : : remoteslot_part);
3688 : : }
3689 : : else
3690 : : {
3691 : 1 : remoteslot_part = ExecCopySlot(remoteslot_part,
3692 : : remoteslot);
3693 : 1 : slot_getallattrs(remoteslot);
3694 : : }
3695 : 1 : MemoryContextSwitchTo(oldctx);
3696 : 1 : apply_handle_insert_internal(edata, partrelinfo_new,
3697 : : remoteslot_part);
3698 : : }
3699 : :
3700 : 11 : EvalPlanQualEnd(&epqstate);
3701 : : }
3702 : 11 : break;
3703 : :
3704 : 0 : default:
3705 [ # # ]: 0 : elog(ERROR, "unrecognized CmdType: %d", (int) operation);
3706 : : break;
3707 : : }
3708 : : }
3709 : :
3710 : : /*
3711 : : * Handle TRUNCATE message.
3712 : : *
3713 : : * TODO: FDW support
3714 : : */
3715 : : static void
3716 : 20 : apply_handle_truncate(StringInfo s)
3717 : : {
3718 : 20 : bool cascade = false;
3719 : 20 : bool restart_seqs = false;
3720 : 20 : List *remote_relids = NIL;
3721 : 20 : List *remote_rels = NIL;
3722 : 20 : List *rels = NIL;
3723 : 20 : List *part_rels = NIL;
3724 : 20 : List *relids = NIL;
3725 : 20 : List *relids_logged = NIL;
3726 : : ListCell *lc;
3727 : 20 : LOCKMODE lockmode = AccessExclusiveLock;
3728 : :
3729 : : /*
3730 : : * Quick return if we are skipping data modification changes or handling
3731 : : * streamed transactions.
3732 : : */
3733 [ + - - + ]: 40 : if (is_skipping_changes() ||
3734 : 20 : handle_streamed_transaction(LOGICAL_REP_MSG_TRUNCATE, s))
3735 : 0 : return;
3736 : :
3737 : 20 : begin_replication_step();
3738 : :
3739 : 20 : remote_relids = logicalrep_read_truncate(s, &cascade, &restart_seqs);
3740 : :
3741 [ + - + + : 49 : foreach(lc, remote_relids)
+ + ]
3742 : : {
3743 : 29 : LogicalRepRelId relid = lfirst_oid(lc);
3744 : : LogicalRepRelMapEntry *rel;
3745 : :
3746 : 29 : rel = logicalrep_rel_open(relid, lockmode);
3747 [ - + ]: 29 : if (!should_apply_changes_for_rel(rel))
3748 : : {
3749 : : /*
3750 : : * The relation can't become interesting in the middle of the
3751 : : * transaction so it's safe to unlock it.
3752 : : */
3753 : 0 : logicalrep_rel_close(rel, lockmode);
3754 : 0 : continue;
3755 : : }
3756 : :
3757 : 29 : remote_rels = lappend(remote_rels, rel);
3758 : 29 : TargetPrivilegesCheck(rel->localrel, ACL_TRUNCATE);
3759 : 29 : rels = lappend(rels, rel->localrel);
3760 : 29 : relids = lappend_oid(relids, rel->localreloid);
3761 [ + + - + : 29 : if (RelationIsLogicallyLogged(rel->localrel))
+ - - + -
- - - + -
+ - ]
3762 : 1 : relids_logged = lappend_oid(relids_logged, rel->localreloid);
3763 : :
3764 : : /*
3765 : : * Truncate partitions if we got a message to truncate a partitioned
3766 : : * table.
3767 : : */
3768 [ + + ]: 29 : if (rel->localrel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE)
3769 : : {
3770 : : ListCell *child;
3771 : 4 : List *children = find_all_inheritors(rel->localreloid,
3772 : : lockmode,
3773 : : NULL);
3774 : :
3775 [ + - + + : 15 : foreach(child, children)
+ + ]
3776 : : {
3777 : 11 : Oid childrelid = lfirst_oid(child);
3778 : : Relation childrel;
3779 : :
3780 [ + + ]: 11 : if (list_member_oid(relids, childrelid))
3781 : 4 : continue;
3782 : :
3783 : : /* find_all_inheritors already got lock */
3784 : 7 : childrel = table_open(childrelid, NoLock);
3785 : :
3786 : : /*
3787 : : * Ignore temp tables of other backends. See similar code in
3788 : : * ExecuteTruncate().
3789 : : */
3790 [ - + - - ]: 7 : if (RELATION_IS_OTHER_TEMP(childrel))
3791 : : {
3792 : 0 : table_close(childrel, lockmode);
3793 : 0 : continue;
3794 : : }
3795 : :
3796 : 7 : TargetPrivilegesCheck(childrel, ACL_TRUNCATE);
3797 : 7 : rels = lappend(rels, childrel);
3798 : 7 : part_rels = lappend(part_rels, childrel);
3799 : 7 : relids = lappend_oid(relids, childrelid);
3800 : : /* Log this relation only if needed for logical decoding */
3801 [ + - - + : 7 : if (RelationIsLogicallyLogged(childrel))
- - - - -
- - - - -
- - ]
3802 : 0 : relids_logged = lappend_oid(relids_logged, childrelid);
3803 : : }
3804 : : }
3805 : : }
3806 : :
3807 : : /*
3808 : : * Even if we used CASCADE on the upstream primary we explicitly default
3809 : : * to replaying changes without further cascading. This might be later
3810 : : * changeable with a user specified option.
3811 : : *
3812 : : * MySubscription->runasowner tells us whether we want to execute
3813 : : * replication actions as the subscription owner; the last argument to
3814 : : * TruncateGuts tells it whether we want to switch to the table owner.
3815 : : * Those are exactly opposite conditions.
3816 : : */
3817 : 20 : ExecuteTruncateGuts(rels,
3818 : : relids,
3819 : : relids_logged,
3820 : : DROP_RESTRICT,
3821 : : restart_seqs,
3822 : 20 : !MySubscription->runasowner);
3823 [ + - + + : 49 : foreach(lc, remote_rels)
+ + ]
3824 : : {
3825 : 29 : LogicalRepRelMapEntry *rel = lfirst(lc);
3826 : :
3827 : 29 : logicalrep_rel_close(rel, NoLock);
3828 : : }
3829 [ + + + + : 27 : foreach(lc, part_rels)
+ + ]
3830 : : {
3831 : 7 : Relation rel = lfirst(lc);
3832 : :
3833 : 7 : table_close(rel, NoLock);
3834 : : }
3835 : :
3836 : 20 : end_replication_step();
3837 : : }
3838 : :
3839 : :
3840 : : /*
3841 : : * Logical replication protocol message dispatcher.
3842 : : */
3843 : : void
3844 : 358330 : apply_dispatch(StringInfo s)
3845 : : {
3846 : 358330 : LogicalRepMsgType action = pq_getmsgbyte(s);
3847 : : LogicalRepMsgType saved_command;
3848 : :
3849 : : /*
3850 : : * Set the current command being applied. Since this function can be
3851 : : * called recursively when applying spooled changes, save the current
3852 : : * command.
3853 : : */
3854 : 358330 : saved_command = remote_ctx.command;
3855 : 358330 : remote_ctx.command = action;
3856 : :
3857 [ + + + + : 358330 : switch (action)
+ + + + +
- + + + +
+ + + + +
- ]
3858 : : {
3859 : 549 : case LOGICAL_REP_MSG_BEGIN:
3860 : 549 : apply_handle_begin(s);
3861 : 549 : break;
3862 : :
3863 : 459 : case LOGICAL_REP_MSG_COMMIT:
3864 : 459 : apply_handle_commit(s);
3865 : 459 : break;
3866 : :
3867 : 206768 : case LOGICAL_REP_MSG_INSERT:
3868 : 206768 : apply_handle_insert(s);
3869 : 206690 : break;
3870 : :
3871 : 66179 : case LOGICAL_REP_MSG_UPDATE:
3872 : 66179 : apply_handle_update(s);
3873 : 66166 : break;
3874 : :
3875 : 81936 : case LOGICAL_REP_MSG_DELETE:
3876 : 81936 : apply_handle_delete(s);
3877 : 81936 : break;
3878 : :
3879 : 20 : case LOGICAL_REP_MSG_TRUNCATE:
3880 : 20 : apply_handle_truncate(s);
3881 : 20 : break;
3882 : :
3883 : 537 : case LOGICAL_REP_MSG_RELATION:
3884 : 537 : apply_handle_relation(s);
3885 : 537 : break;
3886 : :
3887 : 18 : case LOGICAL_REP_MSG_TYPE:
3888 : 18 : apply_handle_type(s);
3889 : 18 : break;
3890 : :
3891 : 7 : case LOGICAL_REP_MSG_ORIGIN:
3892 : 7 : apply_handle_origin(s);
3893 : 7 : break;
3894 : :
3895 : 0 : case LOGICAL_REP_MSG_MESSAGE:
3896 : :
3897 : : /*
3898 : : * Logical replication does not use generic logical messages yet.
3899 : : * Although, it could be used by other applications that use this
3900 : : * output plugin.
3901 : : */
3902 : 0 : break;
3903 : :
3904 : 842 : case LOGICAL_REP_MSG_STREAM_START:
3905 : 842 : apply_handle_stream_start(s);
3906 : 842 : break;
3907 : :
3908 : 841 : case LOGICAL_REP_MSG_STREAM_STOP:
3909 : 841 : apply_handle_stream_stop(s);
3910 : 839 : break;
3911 : :
3912 : 38 : case LOGICAL_REP_MSG_STREAM_ABORT:
3913 : 38 : apply_handle_stream_abort(s);
3914 : 38 : break;
3915 : :
3916 : 61 : case LOGICAL_REP_MSG_STREAM_COMMIT:
3917 : 61 : apply_handle_stream_commit(s);
3918 : 59 : break;
3919 : :
3920 : 17 : case LOGICAL_REP_MSG_BEGIN_PREPARE:
3921 : 17 : apply_handle_begin_prepare(s);
3922 : 17 : break;
3923 : :
3924 : 16 : case LOGICAL_REP_MSG_PREPARE:
3925 : 16 : apply_handle_prepare(s);
3926 : 15 : break;
3927 : :
3928 : 22 : case LOGICAL_REP_MSG_COMMIT_PREPARED:
3929 : 22 : apply_handle_commit_prepared(s);
3930 : 22 : break;
3931 : :
3932 : 5 : case LOGICAL_REP_MSG_ROLLBACK_PREPARED:
3933 : 5 : apply_handle_rollback_prepared(s);
3934 : 5 : break;
3935 : :
3936 : 15 : case LOGICAL_REP_MSG_STREAM_PREPARE:
3937 : 15 : apply_handle_stream_prepare(s);
3938 : 13 : break;
3939 : :
3940 : 0 : default:
3941 [ # # ]: 0 : ereport(ERROR,
3942 : : (errcode(ERRCODE_PROTOCOL_VIOLATION),
3943 : : errmsg("invalid logical replication message type \"??? (%d)\"", action)));
3944 : : }
3945 : :
3946 : : /* Reset the current command */
3947 : 358232 : remote_ctx.command = saved_command;
3948 : 358232 : }
3949 : :
3950 : : /*
3951 : : * Figure out which write/flush positions to report to the walsender process.
3952 : : *
3953 : : * We can't simply report back the last LSN the walsender sent us because the
3954 : : * local transaction might not yet be flushed to disk locally. Instead we
3955 : : * build a list that associates local with remote LSNs for every commit. When
3956 : : * reporting back the flush position to the sender we iterate that list and
3957 : : * check which entries on it are already locally flushed. Those we can report
3958 : : * as having been flushed.
3959 : : *
3960 : : * The have_pending_txes is true if there are outstanding transactions that
3961 : : * need to be flushed.
3962 : : */
3963 : : static void
3964 : 38341 : get_flush_position(XLogRecPtr *write, XLogRecPtr *flush,
3965 : : bool *have_pending_txes)
3966 : : {
3967 : : dlist_mutable_iter iter;
3968 : 38341 : XLogRecPtr local_flush = GetFlushRecPtr(NULL);
3969 : :
3970 : 38341 : *write = InvalidXLogRecPtr;
3971 : 38341 : *flush = InvalidXLogRecPtr;
3972 : :
3973 [ + - + + ]: 38872 : dlist_foreach_modify(iter, &lsn_mapping)
3974 : : {
3975 : 8737 : FlushPosition *pos =
3976 : : dlist_container(FlushPosition, node, iter.cur);
3977 : :
3978 : 8737 : *write = pos->remote_end;
3979 : :
3980 [ + + ]: 8737 : if (pos->local_end <= local_flush)
3981 : : {
3982 : 531 : *flush = pos->remote_end;
3983 : 531 : dlist_delete(iter.cur);
3984 : 531 : pfree(pos);
3985 : : }
3986 : : else
3987 : : {
3988 : : /*
3989 : : * Don't want to uselessly iterate over the rest of the list which
3990 : : * could potentially be long. Instead get the last element and
3991 : : * grab the write position from there.
3992 : : */
3993 : 8206 : pos = dlist_tail_element(FlushPosition, node,
3994 : : &lsn_mapping);
3995 : 8206 : *write = pos->remote_end;
3996 : 8206 : *have_pending_txes = true;
3997 : 8206 : return;
3998 : : }
3999 : : }
4000 : :
4001 : 30135 : *have_pending_txes = !dlist_is_empty(&lsn_mapping);
4002 : : }
4003 : :
4004 : : /*
4005 : : * Store current remote/local lsn pair in the tracking list.
4006 : : */
4007 : : void
4008 : 569 : store_flush_position(XLogRecPtr remote_lsn, XLogRecPtr local_lsn)
4009 : : {
4010 : : FlushPosition *flushpos;
4011 : :
4012 : : /*
4013 : : * Skip for parallel apply workers, because the lsn_mapping is maintained
4014 : : * by the leader apply worker.
4015 : : */
4016 [ + + ]: 569 : if (am_parallel_apply_worker())
4017 : 19 : return;
4018 : :
4019 : : /* Need to do this in permanent context */
4020 : 550 : MemoryContextSwitchTo(ApplyContext);
4021 : :
4022 : : /* Track commit lsn */
4023 : 550 : flushpos = palloc_object(FlushPosition);
4024 : 550 : flushpos->local_end = local_lsn;
4025 : 550 : flushpos->remote_end = remote_lsn;
4026 : :
4027 : 550 : dlist_push_tail(&lsn_mapping, &flushpos->node);
4028 : 550 : MemoryContextSwitchTo(ApplyMessageContext);
4029 : : }
4030 : :
4031 : :
4032 : : /* Update statistics of the worker. */
4033 : : static void
4034 : 221520 : UpdateWorkerStats(XLogRecPtr last_lsn, TimestampTz send_time, bool reply)
4035 : : {
4036 : 221520 : MyLogicalRepWorker->last_lsn = last_lsn;
4037 : 221520 : MyLogicalRepWorker->last_send_time = send_time;
4038 : 221520 : MyLogicalRepWorker->last_recv_time = GetCurrentTimestamp();
4039 [ + + ]: 221520 : if (reply)
4040 : : {
4041 : 4581 : MyLogicalRepWorker->reply_lsn = last_lsn;
4042 : 4581 : MyLogicalRepWorker->reply_time = send_time;
4043 : : }
4044 : 221520 : }
4045 : :
4046 : : /*
4047 : : * Apply main loop.
4048 : : */
4049 : : static void
4050 : 495 : LogicalRepApplyLoop(XLogRecPtr last_received)
4051 : : {
4052 : 495 : TimestampTz last_recv_timestamp = GetCurrentTimestamp();
4053 : 495 : bool ping_sent = false;
4054 : : TimeLineID tli;
4055 : : ErrorContextCallback errcallback;
4056 : 495 : RetainDeadTuplesData rdt_data = {0};
4057 : :
4058 : : /*
4059 : : * Init the ApplyMessageContext which we clean up after each replication
4060 : : * protocol message.
4061 : : */
4062 : 495 : ApplyMessageContext = AllocSetContextCreate(ApplyContext,
4063 : : "ApplyMessageContext",
4064 : : ALLOCSET_DEFAULT_SIZES);
4065 : :
4066 : : /*
4067 : : * This memory context is used for per-stream data when the streaming mode
4068 : : * is enabled. This context is reset on each stream stop.
4069 : : */
4070 : 495 : LogicalStreamingContext = AllocSetContextCreate(ApplyContext,
4071 : : "LogicalStreamingContext",
4072 : : ALLOCSET_DEFAULT_SIZES);
4073 : :
4074 : : /* mark as idle, before starting to loop */
4075 : 495 : pgstat_report_activity(STATE_IDLE, NULL);
4076 : :
4077 : : /*
4078 : : * Push apply error context callback. Fields will be filled while applying
4079 : : * a change.
4080 : : */
4081 : 495 : errcallback.callback = apply_error_callback;
4082 : 495 : errcallback.previous = error_context_stack;
4083 : 495 : error_context_stack = &errcallback;
4084 : 495 : apply_error_context_stack = error_context_stack;
4085 : :
4086 : : /* This outer loop iterates once per wait. */
4087 : : for (;;)
4088 : 32895 : {
4089 : 33390 : pgsocket fd = PGINVALID_SOCKET;
4090 : : int rc;
4091 : : int len;
4092 : 33390 : char *buf = NULL;
4093 : 33390 : bool endofstream = false;
4094 : : long wait_time;
4095 : :
4096 [ - + ]: 33390 : CHECK_FOR_INTERRUPTS();
4097 : :
4098 : 33390 : MemoryContextSwitchTo(ApplyMessageContext);
4099 : :
4100 : 33390 : len = walrcv_receive(LogRepWorkerWalRcvConn, &buf, &fd);
4101 : :
4102 [ + + ]: 33369 : if (len != 0)
4103 : : {
4104 : : /* Loop to process all available data (without blocking). */
4105 : : for (;;)
4106 : : {
4107 [ - + ]: 253421 : CHECK_FOR_INTERRUPTS();
4108 : :
4109 [ + + ]: 253421 : if (len == 0)
4110 : : {
4111 : 31886 : break;
4112 : : }
4113 [ + + ]: 221535 : else if (len < 0)
4114 : : {
4115 [ + - ]: 14 : ereport(LOG,
4116 : : (errmsg("data stream from publisher has ended")));
4117 : 14 : endofstream = true;
4118 : 14 : break;
4119 : : }
4120 : : else
4121 : : {
4122 : : int c;
4123 : : StringInfoData s;
4124 : :
4125 [ - + ]: 221521 : if (ConfigReloadPending)
4126 : : {
4127 : 0 : ConfigReloadPending = false;
4128 : 0 : ProcessConfigFile(PGC_SIGHUP);
4129 : : }
4130 : :
4131 : : /* Reset timeout. */
4132 : 221521 : last_recv_timestamp = GetCurrentTimestamp();
4133 : 221521 : ping_sent = false;
4134 : :
4135 : 221521 : rdt_data.last_recv_time = last_recv_timestamp;
4136 : :
4137 : : /* Ensure we are reading the data into our memory context. */
4138 : 221521 : MemoryContextSwitchTo(ApplyMessageContext);
4139 : :
4140 : 221521 : initReadOnlyStringInfo(&s, buf, len);
4141 : :
4142 : 221521 : c = pq_getmsgbyte(&s);
4143 : :
4144 [ + + ]: 221521 : if (c == PqReplMsg_WALData)
4145 : : {
4146 : : XLogRecPtr start_lsn;
4147 : : XLogRecPtr end_lsn;
4148 : : TimestampTz send_time;
4149 : :
4150 : 205761 : start_lsn = pq_getmsgint64(&s);
4151 : 205761 : end_lsn = pq_getmsgint64(&s);
4152 : 205761 : send_time = pq_getmsgint64(&s);
4153 : :
4154 [ + + ]: 205761 : if (last_received < start_lsn)
4155 : 157782 : last_received = start_lsn;
4156 : :
4157 [ - + ]: 205761 : if (last_received < end_lsn)
4158 : 0 : last_received = end_lsn;
4159 : :
4160 : 205761 : UpdateWorkerStats(last_received, send_time, false);
4161 : :
4162 : 205761 : apply_dispatch(&s);
4163 : :
4164 : 205667 : maybe_advance_nonremovable_xid(&rdt_data, false);
4165 : : }
4166 [ + + ]: 15760 : else if (c == PqReplMsg_Keepalive)
4167 : : {
4168 : : XLogRecPtr end_lsn;
4169 : : TimestampTz timestamp;
4170 : : bool reply_requested;
4171 : :
4172 : 4582 : end_lsn = pq_getmsgint64(&s);
4173 : 4582 : timestamp = pq_getmsgint64(&s);
4174 : 4582 : reply_requested = pq_getmsgbyte(&s);
4175 : :
4176 [ + + ]: 4582 : if (last_received < end_lsn)
4177 : 1184 : last_received = end_lsn;
4178 : :
4179 : 4582 : send_feedback(last_received, reply_requested, false);
4180 : :
4181 : 4581 : maybe_advance_nonremovable_xid(&rdt_data, false);
4182 : :
4183 : 4581 : UpdateWorkerStats(last_received, timestamp, true);
4184 : : }
4185 [ + - ]: 11178 : else if (c == PqReplMsg_PrimaryStatusUpdate)
4186 : : {
4187 : 11178 : rdt_data.remote_lsn = pq_getmsgint64(&s);
4188 : 11178 : rdt_data.remote_oldestxid = FullTransactionIdFromU64((uint64) pq_getmsgint64(&s));
4189 : 11178 : rdt_data.remote_nextxid = FullTransactionIdFromU64((uint64) pq_getmsgint64(&s));
4190 : 11178 : rdt_data.reply_time = pq_getmsgint64(&s);
4191 : :
4192 : : /*
4193 : : * This should never happen, see
4194 : : * ProcessStandbyPSRequestMessage. But if it happens
4195 : : * due to a bug, we don't want to proceed as it can
4196 : : * incorrectly advance oldest_nonremovable_xid.
4197 : : */
4198 [ - + ]: 11178 : if (!XLogRecPtrIsValid(rdt_data.remote_lsn))
4199 [ # # ]: 0 : elog(ERROR, "cannot get the latest WAL position from the publisher");
4200 : :
4201 : 11178 : maybe_advance_nonremovable_xid(&rdt_data, true);
4202 : :
4203 : 11178 : UpdateWorkerStats(last_received, rdt_data.reply_time, false);
4204 : : }
4205 : : /* other message types are purposefully ignored */
4206 : :
4207 : 221426 : MemoryContextReset(ApplyMessageContext);
4208 : : }
4209 : :
4210 : 221426 : len = walrcv_receive(LogRepWorkerWalRcvConn, &buf, &fd);
4211 : : }
4212 : : }
4213 : :
4214 : : /* confirm all writes so far */
4215 : 33273 : send_feedback(last_received, false, false);
4216 : :
4217 : : /* Reset the timestamp if no message was received */
4218 : 33273 : rdt_data.last_recv_time = 0;
4219 : :
4220 : 33273 : maybe_advance_nonremovable_xid(&rdt_data, false);
4221 : :
4222 [ + + + + ]: 33272 : if (!in_remote_transaction && !in_streamed_transaction)
4223 : : {
4224 : : /*
4225 : : * If we didn't get any transactions for a while there might be
4226 : : * unconsumed invalidation messages in the queue, consume them
4227 : : * now.
4228 : : */
4229 : 10881 : AcceptInvalidationMessages();
4230 : 10881 : maybe_reread_subscription();
4231 : :
4232 : : /*
4233 : : * Process any relations that are being synchronized in parallel
4234 : : * and any newly added tables or sequences.
4235 : : */
4236 : 10835 : ProcessSyncingRelations(last_received);
4237 : : }
4238 : :
4239 : : /* Cleanup the memory. */
4240 : 33015 : MemoryContextReset(ApplyMessageContext);
4241 : 33015 : MemoryContextSwitchTo(TopMemoryContext);
4242 : :
4243 : : /* Check if we need to exit the streaming loop. */
4244 [ + + ]: 33015 : if (endofstream)
4245 : 14 : break;
4246 : :
4247 : : /*
4248 : : * Wait for more data or latch. If we have unflushed transactions,
4249 : : * wake up after WalWriterDelay to see if they've been flushed yet (in
4250 : : * which case we should send a feedback message). Otherwise, there's
4251 : : * no particular urgency about waking up unless we get data or a
4252 : : * signal.
4253 : : */
4254 [ + + ]: 33001 : if (!dlist_is_empty(&lsn_mapping))
4255 : 5175 : wait_time = WalWriterDelay;
4256 : : else
4257 : 27826 : wait_time = NAPTIME_PER_CYCLE;
4258 : :
4259 : : /*
4260 : : * Ensure to wake up when it's possible to advance the non-removable
4261 : : * transaction ID, or when the retention duration may have exceeded
4262 : : * max_retention_duration.
4263 : : */
4264 [ + + ]: 33001 : if (MySubscription->retentionactive)
4265 : : {
4266 [ + + ]: 5371 : if (rdt_data.phase == RDT_GET_CANDIDATE_XID &&
4267 [ + - ]: 264 : rdt_data.xid_advance_interval)
4268 : 264 : wait_time = Min(wait_time, rdt_data.xid_advance_interval);
4269 [ + + ]: 5107 : else if (MySubscription->maxretention > 0)
4270 : 1 : wait_time = Min(wait_time, MySubscription->maxretention);
4271 : : }
4272 : :
4273 : 33001 : rc = WaitLatchOrSocket(MyLatch,
4274 : : WL_SOCKET_READABLE | WL_LATCH_SET |
4275 : : WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
4276 : : fd, wait_time,
4277 : : WAIT_EVENT_LOGICAL_APPLY_MAIN);
4278 : :
4279 [ + + ]: 33001 : if (rc & WL_LATCH_SET)
4280 : : {
4281 : 823 : ResetLatch(MyLatch);
4282 [ + + ]: 823 : CHECK_FOR_INTERRUPTS();
4283 : : }
4284 : :
4285 [ + + ]: 32895 : if (ConfigReloadPending)
4286 : : {
4287 : 12 : ConfigReloadPending = false;
4288 : 12 : ProcessConfigFile(PGC_SIGHUP);
4289 : : }
4290 : :
4291 [ + + ]: 32895 : if (rc & WL_TIMEOUT)
4292 : : {
4293 : : /*
4294 : : * We didn't receive anything new. If we haven't heard anything
4295 : : * from the server for more than wal_receiver_timeout / 2, ping
4296 : : * the server. Also, if it's been longer than
4297 : : * wal_receiver_status_interval since the last update we sent,
4298 : : * send a status update to the primary anyway, to report any
4299 : : * progress in applying WAL.
4300 : : */
4301 : 471 : bool requestReply = false;
4302 : :
4303 : : /*
4304 : : * Check if time since last receive from primary has reached the
4305 : : * configured limit.
4306 : : */
4307 [ + - ]: 471 : if (wal_receiver_timeout > 0)
4308 : : {
4309 : 471 : TimestampTz now = GetCurrentTimestamp();
4310 : : TimestampTz timeout;
4311 : :
4312 : 471 : timeout =
4313 : 471 : TimestampTzPlusMilliseconds(last_recv_timestamp,
4314 : : wal_receiver_timeout);
4315 : :
4316 [ - + ]: 471 : if (now >= timeout)
4317 [ # # ]: 0 : ereport(ERROR,
4318 : : (errcode(ERRCODE_CONNECTION_FAILURE),
4319 : : errmsg("terminating logical replication worker due to timeout")));
4320 : :
4321 : : /* Check to see if it's time for a ping. */
4322 [ + - ]: 471 : if (!ping_sent)
4323 : : {
4324 : 471 : timeout = TimestampTzPlusMilliseconds(last_recv_timestamp,
4325 : : (wal_receiver_timeout / 2));
4326 [ - + ]: 471 : if (now >= timeout)
4327 : : {
4328 : 0 : requestReply = true;
4329 : 0 : ping_sent = true;
4330 : : }
4331 : : }
4332 : : }
4333 : :
4334 : 471 : send_feedback(last_received, requestReply, requestReply);
4335 : :
4336 : 471 : maybe_advance_nonremovable_xid(&rdt_data, false);
4337 : :
4338 : : /*
4339 : : * Force reporting to ensure long idle periods don't lead to
4340 : : * arbitrarily delayed stats. Stats can only be reported outside
4341 : : * of (implicit or explicit) transactions. That shouldn't lead to
4342 : : * stats being delayed for long, because transactions are either
4343 : : * sent as a whole on commit or streamed. Streamed transactions
4344 : : * are spilled to disk and applied on commit.
4345 : : */
4346 [ + - ]: 471 : if (!IsTransactionState())
4347 : 471 : pgstat_report_stat(true);
4348 : : }
4349 : : }
4350 : :
4351 : : /* Pop the error context stack */
4352 : 14 : error_context_stack = errcallback.previous;
4353 : 14 : apply_error_context_stack = error_context_stack;
4354 : :
4355 : : /* All done */
4356 : 14 : walrcv_endstreaming(LogRepWorkerWalRcvConn, &tli);
4357 : 0 : }
4358 : :
4359 : : /*
4360 : : * Send a Standby Status Update message to server.
4361 : : *
4362 : : * 'recvpos' is the latest LSN we've received data to, force is set if we need
4363 : : * to send a response to avoid timeouts.
4364 : : */
4365 : : static void
4366 : 38326 : send_feedback(XLogRecPtr recvpos, bool force, bool requestReply)
4367 : : {
4368 : : static StringInfo reply_message = NULL;
4369 : : static TimestampTz send_time = 0;
4370 : :
4371 : : static XLogRecPtr last_recvpos = InvalidXLogRecPtr;
4372 : : static XLogRecPtr last_writepos = InvalidXLogRecPtr;
4373 : :
4374 : : XLogRecPtr writepos;
4375 : : XLogRecPtr flushpos;
4376 : : TimestampTz now;
4377 : : bool have_pending_txes;
4378 : :
4379 : : /*
4380 : : * If the user doesn't want status to be reported to the publisher, be
4381 : : * sure to exit before doing anything at all.
4382 : : */
4383 [ + + - + ]: 38326 : if (!force && wal_receiver_status_interval <= 0)
4384 : 16399 : return;
4385 : :
4386 : : /* It's legal to not pass a recvpos */
4387 [ - + ]: 38326 : if (recvpos < last_recvpos)
4388 : 0 : recvpos = last_recvpos;
4389 : :
4390 : 38326 : get_flush_position(&writepos, &flushpos, &have_pending_txes);
4391 : :
4392 : : /*
4393 : : * No outstanding transactions to flush, we can report the latest received
4394 : : * position. This is important for synchronous replication.
4395 : : */
4396 [ + + ]: 38326 : if (!have_pending_txes)
4397 : 30124 : flushpos = writepos = recvpos;
4398 : :
4399 [ - + ]: 38326 : if (writepos < last_writepos)
4400 : 0 : writepos = last_writepos;
4401 : :
4402 [ + + ]: 38326 : if (flushpos < last_flushpos)
4403 : 8150 : flushpos = last_flushpos;
4404 : :
4405 : 38326 : now = GetCurrentTimestamp();
4406 : :
4407 : : /* if we've already reported everything we're good */
4408 [ + + ]: 38326 : if (!force &&
4409 [ + + ]: 35787 : writepos == last_writepos &&
4410 [ + + ]: 16733 : flushpos == last_flushpos &&
4411 [ + + ]: 16506 : !TimestampDifferenceExceeds(send_time, now,
4412 : : wal_receiver_status_interval * 1000))
4413 : 16399 : return;
4414 : 21927 : send_time = now;
4415 : :
4416 [ + + ]: 21927 : if (!reply_message)
4417 : : {
4418 : 495 : MemoryContext oldctx = MemoryContextSwitchTo(ApplyContext);
4419 : :
4420 : 495 : reply_message = makeStringInfo();
4421 : 495 : MemoryContextSwitchTo(oldctx);
4422 : : }
4423 : : else
4424 : 21432 : resetStringInfo(reply_message);
4425 : :
4426 : 21927 : pq_sendbyte(reply_message, PqReplMsg_StandbyStatusUpdate);
4427 : 21927 : pq_sendint64(reply_message, recvpos); /* write */
4428 : 21927 : pq_sendint64(reply_message, flushpos); /* flush */
4429 : 21927 : pq_sendint64(reply_message, writepos); /* apply */
4430 : 21927 : pq_sendint64(reply_message, now); /* sendTime */
4431 : 21927 : pq_sendbyte(reply_message, requestReply); /* replyRequested */
4432 : :
4433 [ + + ]: 21927 : elog(DEBUG2, "sending feedback (force %d) to recv %X/%08X, write %X/%08X, flush %X/%08X",
4434 : : force,
4435 : : LSN_FORMAT_ARGS(recvpos),
4436 : : LSN_FORMAT_ARGS(writepos),
4437 : : LSN_FORMAT_ARGS(flushpos));
4438 : :
4439 : 21927 : walrcv_send(LogRepWorkerWalRcvConn,
4440 : : reply_message->data, reply_message->len);
4441 : :
4442 [ + + ]: 21926 : if (recvpos > last_recvpos)
4443 : 19054 : last_recvpos = recvpos;
4444 [ + + ]: 21926 : if (writepos > last_writepos)
4445 : 19059 : last_writepos = writepos;
4446 [ + + ]: 21926 : if (flushpos > last_flushpos)
4447 : 18895 : last_flushpos = flushpos;
4448 : : }
4449 : :
4450 : : /*
4451 : : * Attempt to advance the non-removable transaction ID.
4452 : : *
4453 : : * See comments atop worker.c for details.
4454 : : */
4455 : : static void
4456 : 255170 : maybe_advance_nonremovable_xid(RetainDeadTuplesData *rdt_data,
4457 : : bool status_received)
4458 : : {
4459 [ + + ]: 255170 : if (!can_advance_nonremovable_xid(rdt_data))
4460 : 238233 : return;
4461 : :
4462 : 16937 : process_rdt_phase_transition(rdt_data, status_received);
4463 : : }
4464 : :
4465 : : /*
4466 : : * Preliminary check to determine if advancing the non-removable transaction ID
4467 : : * is allowed.
4468 : : */
4469 : : static bool
4470 : 255170 : can_advance_nonremovable_xid(RetainDeadTuplesData *rdt_data)
4471 : : {
4472 : : /*
4473 : : * It is sufficient to manage non-removable transaction ID for a
4474 : : * subscription by the main apply worker to detect update_deleted reliably
4475 : : * even for table sync or parallel apply workers.
4476 : : */
4477 [ + + ]: 255170 : if (!am_leader_apply_worker())
4478 : 395 : return false;
4479 : :
4480 : : /* No need to advance if retaining dead tuples is not required */
4481 [ + + ]: 254775 : if (!MySubscription->retaindeadtuples)
4482 : 237838 : return false;
4483 : :
4484 : 16937 : return true;
4485 : : }
4486 : :
4487 : : /*
4488 : : * Process phase transitions during the non-removable transaction ID
4489 : : * advancement. See comments atop worker.c for details of the transition.
4490 : : */
4491 : : static void
4492 : 28250 : process_rdt_phase_transition(RetainDeadTuplesData *rdt_data,
4493 : : bool status_received)
4494 : : {
4495 [ + + + + : 28250 : switch (rdt_data->phase)
+ + - ]
4496 : : {
4497 : 659 : case RDT_GET_CANDIDATE_XID:
4498 : 659 : get_candidate_xid(rdt_data);
4499 : 659 : break;
4500 : 11179 : case RDT_REQUEST_PUBLISHER_STATUS:
4501 : 11179 : request_publisher_status(rdt_data);
4502 : 11179 : break;
4503 : 16285 : case RDT_WAIT_FOR_PUBLISHER_STATUS:
4504 : 16285 : wait_for_publisher_status(rdt_data, status_received);
4505 : 16285 : break;
4506 : 125 : case RDT_WAIT_FOR_LOCAL_FLUSH:
4507 : 125 : wait_for_local_flush(rdt_data);
4508 : 125 : break;
4509 : 1 : case RDT_STOP_CONFLICT_INFO_RETENTION:
4510 : 1 : stop_conflict_info_retention(rdt_data);
4511 : 1 : break;
4512 : 1 : case RDT_RESUME_CONFLICT_INFO_RETENTION:
4513 : 1 : resume_conflict_info_retention(rdt_data);
4514 : 0 : break;
4515 : : }
4516 : 28249 : }
4517 : :
4518 : : /*
4519 : : * Workhorse for the RDT_GET_CANDIDATE_XID phase.
4520 : : */
4521 : : static void
4522 : 659 : get_candidate_xid(RetainDeadTuplesData *rdt_data)
4523 : : {
4524 : : TransactionId oldest_running_xid;
4525 : : TimestampTz now;
4526 : :
4527 : : /*
4528 : : * Use last_recv_time when applying changes in the loop to avoid
4529 : : * unnecessary system time retrieval. If last_recv_time is not available,
4530 : : * obtain the current timestamp.
4531 : : */
4532 [ + + ]: 659 : now = rdt_data->last_recv_time ? rdt_data->last_recv_time : GetCurrentTimestamp();
4533 : :
4534 : : /*
4535 : : * Compute the candidate_xid and request the publisher status at most once
4536 : : * per xid_advance_interval. Refer to adjust_xid_advance_interval() for
4537 : : * details on how this value is dynamically adjusted. This is to avoid
4538 : : * using CPU and network resources without making much progress.
4539 : : */
4540 [ + + ]: 659 : if (!TimestampDifferenceExceeds(rdt_data->candidate_xid_time, now,
4541 : : rdt_data->xid_advance_interval))
4542 : 511 : return;
4543 : :
4544 : : /*
4545 : : * Immediately update the timer, even if the function returns later
4546 : : * without setting candidate_xid due to inactivity on the subscriber. This
4547 : : * avoids frequent calls to GetOldestActiveTransactionId.
4548 : : */
4549 : 148 : rdt_data->candidate_xid_time = now;
4550 : :
4551 : : /*
4552 : : * Consider transactions in the current database, as only dead tuples from
4553 : : * this database are required for conflict detection.
4554 : : */
4555 : 148 : oldest_running_xid = GetOldestActiveTransactionId(false, false);
4556 : :
4557 : : /*
4558 : : * Oldest active transaction ID (oldest_running_xid) can't be behind any
4559 : : * of its previously computed value.
4560 : : */
4561 : : Assert(TransactionIdPrecedesOrEquals(MyLogicalRepWorker->oldest_nonremovable_xid,
4562 : : oldest_running_xid));
4563 : :
4564 : : /* Return if the oldest_nonremovable_xid cannot be advanced */
4565 [ + + ]: 148 : if (TransactionIdEquals(MyLogicalRepWorker->oldest_nonremovable_xid,
4566 : : oldest_running_xid))
4567 : : {
4568 : 79 : adjust_xid_advance_interval(rdt_data, false);
4569 : 79 : return;
4570 : : }
4571 : :
4572 : 69 : adjust_xid_advance_interval(rdt_data, true);
4573 : :
4574 : 69 : rdt_data->candidate_xid = oldest_running_xid;
4575 : 69 : rdt_data->phase = RDT_REQUEST_PUBLISHER_STATUS;
4576 : :
4577 : : /* process the next phase */
4578 : 69 : process_rdt_phase_transition(rdt_data, false);
4579 : : }
4580 : :
4581 : : /*
4582 : : * Workhorse for the RDT_REQUEST_PUBLISHER_STATUS phase.
4583 : : */
4584 : : static void
4585 : 11179 : request_publisher_status(RetainDeadTuplesData *rdt_data)
4586 : : {
4587 : : static StringInfo request_message = NULL;
4588 : :
4589 [ + + ]: 11179 : if (!request_message)
4590 : : {
4591 : 14 : MemoryContext oldctx = MemoryContextSwitchTo(ApplyContext);
4592 : :
4593 : 14 : request_message = makeStringInfo();
4594 : 14 : MemoryContextSwitchTo(oldctx);
4595 : : }
4596 : : else
4597 : 11165 : resetStringInfo(request_message);
4598 : :
4599 : : /*
4600 : : * Send the current time to update the remote walsender's latest reply
4601 : : * message received time.
4602 : : */
4603 : 11179 : pq_sendbyte(request_message, PqReplMsg_PrimaryStatusRequest);
4604 : 11179 : pq_sendint64(request_message, GetCurrentTimestamp());
4605 : :
4606 [ + + ]: 11179 : elog(DEBUG2, "sending publisher status request message");
4607 : :
4608 : : /* Send a request for the publisher status */
4609 : 11179 : walrcv_send(LogRepWorkerWalRcvConn,
4610 : : request_message->data, request_message->len);
4611 : :
4612 : 11179 : rdt_data->phase = RDT_WAIT_FOR_PUBLISHER_STATUS;
4613 : :
4614 : : /*
4615 : : * Skip calling maybe_advance_nonremovable_xid() since further transition
4616 : : * is possible only once we receive the publisher status message.
4617 : : */
4618 : 11179 : }
4619 : :
4620 : : /*
4621 : : * Workhorse for the RDT_WAIT_FOR_PUBLISHER_STATUS phase.
4622 : : */
4623 : : static void
4624 : 16285 : wait_for_publisher_status(RetainDeadTuplesData *rdt_data,
4625 : : bool status_received)
4626 : : {
4627 : : /*
4628 : : * Return if we have requested but not yet received the publisher status.
4629 : : */
4630 [ + + ]: 16285 : if (!status_received)
4631 : 5107 : return;
4632 : :
4633 : : /*
4634 : : * We don't need to maintain oldest_nonremovable_xid if we decide to stop
4635 : : * retaining conflict information for this worker.
4636 : : */
4637 [ - + ]: 11178 : if (should_stop_conflict_info_retention(rdt_data))
4638 : : {
4639 : 0 : rdt_data->phase = RDT_STOP_CONFLICT_INFO_RETENTION;
4640 : 0 : return;
4641 : : }
4642 : :
4643 [ + + ]: 11178 : if (!FullTransactionIdIsValid(rdt_data->remote_wait_for))
4644 : 68 : rdt_data->remote_wait_for = rdt_data->remote_nextxid;
4645 : :
4646 : : /*
4647 : : * Check if all remote concurrent transactions that were active at the
4648 : : * first status request have now completed. If completed, proceed to the
4649 : : * next phase; otherwise, continue checking the publisher status until
4650 : : * these transactions finish.
4651 : : *
4652 : : * It's possible that transactions in the commit phase during the last
4653 : : * cycle have now finished committing, but remote_oldestxid remains older
4654 : : * than remote_wait_for. This can happen if some old transaction came in
4655 : : * the commit phase when we requested status in this cycle. We do not
4656 : : * handle this case explicitly as it's rare and the benefit doesn't
4657 : : * justify the required complexity. Tracking would require either caching
4658 : : * all xids at the publisher or sending them to subscribers. The condition
4659 : : * will resolve naturally once the remaining transactions are finished.
4660 : : *
4661 : : * Directly advancing the non-removable transaction ID is possible if
4662 : : * there are no activities on the publisher since the last advancement
4663 : : * cycle. However, it requires maintaining two fields, last_remote_nextxid
4664 : : * and last_remote_lsn, within the structure for comparison with the
4665 : : * current cycle's values. Considering the minimal cost of continuing in
4666 : : * RDT_WAIT_FOR_LOCAL_FLUSH without awaiting changes, we opted not to
4667 : : * advance the transaction ID here.
4668 : : */
4669 [ + + ]: 11178 : if (FullTransactionIdPrecedesOrEquals(rdt_data->remote_wait_for,
4670 : : rdt_data->remote_oldestxid))
4671 : 68 : rdt_data->phase = RDT_WAIT_FOR_LOCAL_FLUSH;
4672 : : else
4673 : 11110 : rdt_data->phase = RDT_REQUEST_PUBLISHER_STATUS;
4674 : :
4675 : : /* process the next phase */
4676 : 11178 : process_rdt_phase_transition(rdt_data, false);
4677 : : }
4678 : :
4679 : : /*
4680 : : * Workhorse for the RDT_WAIT_FOR_LOCAL_FLUSH phase.
4681 : : */
4682 : : static void
4683 : 125 : wait_for_local_flush(RetainDeadTuplesData *rdt_data)
4684 : : {
4685 : : Assert(XLogRecPtrIsValid(rdt_data->remote_lsn) &&
4686 : : TransactionIdIsValid(rdt_data->candidate_xid));
4687 : :
4688 : : /*
4689 : : * We expect the publisher and subscriber clocks to be in sync using time
4690 : : * sync service like NTP. Otherwise, we will advance this worker's
4691 : : * oldest_nonremovable_xid prematurely, leading to the removal of rows
4692 : : * required to detect update_deleted reliably. This check primarily
4693 : : * addresses scenarios where the publisher's clock falls behind; if the
4694 : : * publisher's clock is ahead, subsequent transactions will naturally bear
4695 : : * later commit timestamps, conforming to the design outlined atop
4696 : : * worker.c.
4697 : : *
4698 : : * XXX Consider waiting for the publisher's clock to catch up with the
4699 : : * subscriber's before proceeding to the next phase.
4700 : : */
4701 [ - + ]: 125 : if (TimestampDifferenceExceeds(rdt_data->reply_time,
4702 : : rdt_data->candidate_xid_time, 0))
4703 [ # # ]: 0 : ereport(ERROR,
4704 : : errmsg_internal("oldest_nonremovable_xid transaction ID could be advanced prematurely"),
4705 : : errdetail_internal("The clock on the publisher is behind that of the subscriber."));
4706 : :
4707 : : /*
4708 : : * Do not attempt to advance the non-removable transaction ID when table
4709 : : * sync is in progress. During this time, changes from a single
4710 : : * transaction may be applied by multiple table sync workers corresponding
4711 : : * to the target tables. So, it's necessary for all table sync workers to
4712 : : * apply and flush the corresponding changes before advancing the
4713 : : * transaction ID, otherwise, dead tuples that are still needed for
4714 : : * conflict detection in table sync workers could be removed prematurely.
4715 : : * However, confirming the apply and flush progress across all table sync
4716 : : * workers is complex and not worth the effort, so we simply return if not
4717 : : * all tables are in the READY state.
4718 : : *
4719 : : * Advancing the transaction ID is necessary even when no tables are
4720 : : * currently subscribed, to avoid retaining dead tuples unnecessarily.
4721 : : * While it might seem safe to skip all phases and directly assign
4722 : : * candidate_xid to oldest_nonremovable_xid during the
4723 : : * RDT_GET_CANDIDATE_XID phase in such cases, this is unsafe. If users
4724 : : * concurrently add tables to the subscription, the apply worker may not
4725 : : * process invalidations in time. Consequently,
4726 : : * HasSubscriptionTablesCached() might miss the new tables, leading to
4727 : : * premature advancement of oldest_nonremovable_xid.
4728 : : *
4729 : : * Performing the check during RDT_WAIT_FOR_LOCAL_FLUSH is safe, as
4730 : : * invalidations are guaranteed to be processed before applying changes
4731 : : * from newly added tables while waiting for the local flush to reach
4732 : : * remote_lsn.
4733 : : *
4734 : : * Additionally, even if we check for subscription tables during
4735 : : * RDT_GET_CANDIDATE_XID, they might be dropped before reaching
4736 : : * RDT_WAIT_FOR_LOCAL_FLUSH. Therefore, it's still necessary to verify
4737 : : * subscription tables at this stage to prevent unnecessary tuple
4738 : : * retention.
4739 : : */
4740 [ + + + + ]: 125 : if (HasSubscriptionTablesCached() && !AllTablesyncsReady())
4741 : : {
4742 : : TimestampTz now;
4743 : :
4744 : 8 : now = rdt_data->last_recv_time
4745 [ + + ]: 4 : ? rdt_data->last_recv_time : GetCurrentTimestamp();
4746 : :
4747 : : /*
4748 : : * Record the time spent waiting for table sync, it is needed for the
4749 : : * timeout check in should_stop_conflict_info_retention().
4750 : : */
4751 : 4 : rdt_data->table_sync_wait_time =
4752 : 4 : TimestampDifferenceMilliseconds(rdt_data->candidate_xid_time, now);
4753 : :
4754 : 4 : return;
4755 : : }
4756 : :
4757 : : /*
4758 : : * We don't need to maintain oldest_nonremovable_xid if we decide to stop
4759 : : * retaining conflict information for this worker.
4760 : : */
4761 [ + + ]: 121 : if (should_stop_conflict_info_retention(rdt_data))
4762 : : {
4763 : 1 : rdt_data->phase = RDT_STOP_CONFLICT_INFO_RETENTION;
4764 : 1 : return;
4765 : : }
4766 : :
4767 : : /*
4768 : : * Update and check the remote flush position if we are applying changes
4769 : : * in a loop. This is done at most once per WalWriterDelay to avoid
4770 : : * performing costly operations in get_flush_position() too frequently
4771 : : * during change application.
4772 : : */
4773 [ + + + + : 153 : if (last_flushpos < rdt_data->remote_lsn && rdt_data->last_recv_time &&
+ + ]
4774 : 33 : TimestampDifferenceExceeds(rdt_data->flushpos_update_time,
4775 : : rdt_data->last_recv_time, WalWriterDelay))
4776 : : {
4777 : : XLogRecPtr writepos;
4778 : : XLogRecPtr flushpos;
4779 : : bool have_pending_txes;
4780 : :
4781 : : /* Fetch the latest remote flush position */
4782 : 15 : get_flush_position(&writepos, &flushpos, &have_pending_txes);
4783 : :
4784 [ - + ]: 15 : if (flushpos > last_flushpos)
4785 : 0 : last_flushpos = flushpos;
4786 : :
4787 : 15 : rdt_data->flushpos_update_time = rdt_data->last_recv_time;
4788 : : }
4789 : :
4790 : : /* Return to wait for the changes to be applied */
4791 [ + + ]: 120 : if (last_flushpos < rdt_data->remote_lsn)
4792 : 53 : return;
4793 : :
4794 : : /*
4795 : : * Reaching this point implies should_stop_conflict_info_retention()
4796 : : * returned false earlier, meaning that the most recent duration for
4797 : : * advancing the non-removable transaction ID is within the
4798 : : * max_retention_duration or max_retention_duration is set to 0.
4799 : : *
4800 : : * Therefore, if conflict info retention was previously stopped due to a
4801 : : * timeout, it is now safe to resume retention.
4802 : : */
4803 [ + + ]: 67 : if (!MySubscription->retentionactive)
4804 : : {
4805 : 1 : rdt_data->phase = RDT_RESUME_CONFLICT_INFO_RETENTION;
4806 : 1 : return;
4807 : : }
4808 : :
4809 : : /*
4810 : : * Reaching here means the remote WAL position has been received, and all
4811 : : * transactions up to that position on the publisher have been applied and
4812 : : * flushed locally. So, we can advance the non-removable transaction ID.
4813 : : */
4814 : 66 : SpinLockAcquire(&MyLogicalRepWorker->relmutex);
4815 : 66 : MyLogicalRepWorker->oldest_nonremovable_xid = rdt_data->candidate_xid;
4816 : 66 : SpinLockRelease(&MyLogicalRepWorker->relmutex);
4817 : :
4818 [ + + ]: 66 : elog(DEBUG2, "confirmed flush up to remote lsn %X/%08X: new oldest_nonremovable_xid %u",
4819 : : LSN_FORMAT_ARGS(rdt_data->remote_lsn),
4820 : : rdt_data->candidate_xid);
4821 : :
4822 : : /* Notify launcher to update the xmin of the conflict slot */
4823 : 66 : ApplyLauncherWakeup();
4824 : :
4825 : 66 : reset_retention_data_fields(rdt_data);
4826 : :
4827 : : /* process the next phase */
4828 : 66 : process_rdt_phase_transition(rdt_data, false);
4829 : : }
4830 : :
4831 : : /*
4832 : : * Check whether conflict information retention should be stopped due to
4833 : : * exceeding the maximum wait time (max_retention_duration).
4834 : : *
4835 : : * If retention should be stopped, return true. Otherwise, return false.
4836 : : */
4837 : : static bool
4838 : 11299 : should_stop_conflict_info_retention(RetainDeadTuplesData *rdt_data)
4839 : : {
4840 : : TimestampTz now;
4841 : :
4842 : : Assert(TransactionIdIsValid(rdt_data->candidate_xid));
4843 : : Assert(rdt_data->phase == RDT_WAIT_FOR_PUBLISHER_STATUS ||
4844 : : rdt_data->phase == RDT_WAIT_FOR_LOCAL_FLUSH);
4845 : :
4846 [ + + ]: 11299 : if (!MySubscription->maxretention)
4847 : 11298 : return false;
4848 : :
4849 : : /*
4850 : : * Use last_recv_time when applying changes in the loop to avoid
4851 : : * unnecessary system time retrieval. If last_recv_time is not available,
4852 : : * obtain the current timestamp.
4853 : : */
4854 [ - + ]: 1 : now = rdt_data->last_recv_time ? rdt_data->last_recv_time : GetCurrentTimestamp();
4855 : :
4856 : : /*
4857 : : * Return early if the wait time has not exceeded the configured maximum
4858 : : * (max_retention_duration). Time spent waiting for table synchronization
4859 : : * is excluded from this calculation, as it occurs infrequently.
4860 : : */
4861 [ - + ]: 1 : if (!TimestampDifferenceExceeds(rdt_data->candidate_xid_time, now,
4862 : 1 : MySubscription->maxretention +
4863 : 1 : rdt_data->table_sync_wait_time))
4864 : 0 : return false;
4865 : :
4866 : 1 : return true;
4867 : : }
4868 : :
4869 : : /*
4870 : : * Workhorse for the RDT_STOP_CONFLICT_INFO_RETENTION phase.
4871 : : */
4872 : : static void
4873 : 1 : stop_conflict_info_retention(RetainDeadTuplesData *rdt_data)
4874 : : {
4875 : : /* Stop retention if not yet */
4876 [ + - ]: 1 : if (MySubscription->retentionactive)
4877 : : {
4878 : : /*
4879 : : * If the retention status cannot be updated (e.g., due to active
4880 : : * transaction), skip further processing to avoid inconsistent
4881 : : * retention behavior.
4882 : : */
4883 [ - + ]: 1 : if (!update_retention_status(false))
4884 : 0 : return;
4885 : :
4886 : 1 : SpinLockAcquire(&MyLogicalRepWorker->relmutex);
4887 : 1 : MyLogicalRepWorker->oldest_nonremovable_xid = InvalidTransactionId;
4888 : 1 : SpinLockRelease(&MyLogicalRepWorker->relmutex);
4889 : :
4890 [ + - ]: 1 : ereport(LOG,
4891 : : errmsg("logical replication worker for subscription \"%s\" has stopped retaining the information for detecting conflicts",
4892 : : MySubscription->name),
4893 : : errdetail("Retention is stopped because the apply process has not caught up with the publisher within the configured max_retention_duration."));
4894 : : }
4895 : :
4896 : : Assert(!TransactionIdIsValid(MyLogicalRepWorker->oldest_nonremovable_xid));
4897 : :
4898 : : /*
4899 : : * If retention has been stopped, reset to the initial phase to retry
4900 : : * resuming retention. This reset is required to recalculate the current
4901 : : * wait time and resume retention if the time falls within
4902 : : * max_retention_duration.
4903 : : */
4904 : 1 : reset_retention_data_fields(rdt_data);
4905 : : }
4906 : :
4907 : : /*
4908 : : * Workhorse for the RDT_RESUME_CONFLICT_INFO_RETENTION phase.
4909 : : */
4910 : : static void
4911 : 1 : resume_conflict_info_retention(RetainDeadTuplesData *rdt_data)
4912 : : {
4913 : : /* We can't resume retention without updating retention status. */
4914 [ - + ]: 1 : if (!update_retention_status(true))
4915 : 0 : return;
4916 : :
4917 [ + - - + ]: 1 : ereport(LOG,
4918 : : errmsg("logical replication worker for subscription \"%s\" will resume retaining the information for detecting conflicts",
4919 : : MySubscription->name),
4920 : : MySubscription->maxretention
4921 : : ? errdetail("Retention is re-enabled because the apply process has caught up with the publisher within the configured max_retention_duration.")
4922 : : : errdetail("Retention is re-enabled because max_retention_duration has been set to unlimited."));
4923 : :
4924 : : /*
4925 : : * Restart the worker to let the launcher initialize
4926 : : * oldest_nonremovable_xid at startup.
4927 : : *
4928 : : * While it's technically possible to derive this value on-the-fly using
4929 : : * the conflict detection slot's xmin, doing so risks a race condition:
4930 : : * the launcher might clean slot.xmin just after retention resumes. This
4931 : : * would make oldest_nonremovable_xid unreliable, especially during xid
4932 : : * wraparound.
4933 : : *
4934 : : * Although this can be prevented by introducing heavy weight locking, the
4935 : : * complexity it will bring doesn't seem worthwhile given how rarely
4936 : : * retention is resumed.
4937 : : */
4938 : 1 : apply_worker_exit();
4939 : : }
4940 : :
4941 : : /*
4942 : : * Updates pg_subscription.subretentionactive to the given value within a
4943 : : * new transaction.
4944 : : *
4945 : : * If already inside an active transaction, skips the update and returns
4946 : : * false.
4947 : : *
4948 : : * Returns true if the update is successfully performed.
4949 : : */
4950 : : static bool
4951 : 2 : update_retention_status(bool active)
4952 : : {
4953 : : /*
4954 : : * Do not update the catalog during an active transaction. The transaction
4955 : : * may be started during change application, leading to a possible
4956 : : * rollback of catalog updates if the application fails subsequently.
4957 : : */
4958 [ - + ]: 2 : if (IsTransactionState())
4959 : 0 : return false;
4960 : :
4961 : 2 : StartTransactionCommand();
4962 : :
4963 : : /*
4964 : : * Updating pg_subscription might involve TOAST table access, so ensure we
4965 : : * have a valid snapshot.
4966 : : */
4967 : 2 : PushActiveSnapshot(GetTransactionSnapshot());
4968 : :
4969 : : /* Update pg_subscription.subretentionactive */
4970 : 2 : UpdateDeadTupleRetentionStatus(MySubscription->oid, active);
4971 : :
4972 : 2 : PopActiveSnapshot();
4973 : 2 : CommitTransactionCommand();
4974 : :
4975 : : /* Notify launcher to update the conflict slot */
4976 : 2 : ApplyLauncherWakeup();
4977 : :
4978 : 2 : MySubscription->retentionactive = active;
4979 : :
4980 : 2 : return true;
4981 : : }
4982 : :
4983 : : /*
4984 : : * Reset all data fields of RetainDeadTuplesData except those used to
4985 : : * determine the timing for the next round of transaction ID advancement. We
4986 : : * can even use flushpos_update_time in the next round to decide whether to get
4987 : : * the latest flush position.
4988 : : */
4989 : : static void
4990 : 67 : reset_retention_data_fields(RetainDeadTuplesData *rdt_data)
4991 : : {
4992 : 67 : rdt_data->phase = RDT_GET_CANDIDATE_XID;
4993 : 67 : rdt_data->remote_lsn = InvalidXLogRecPtr;
4994 : 67 : rdt_data->remote_oldestxid = InvalidFullTransactionId;
4995 : 67 : rdt_data->remote_nextxid = InvalidFullTransactionId;
4996 : 67 : rdt_data->reply_time = 0;
4997 : 67 : rdt_data->remote_wait_for = InvalidFullTransactionId;
4998 : 67 : rdt_data->candidate_xid = InvalidTransactionId;
4999 : 67 : rdt_data->table_sync_wait_time = 0;
5000 : 67 : }
5001 : :
5002 : : /*
5003 : : * Adjust the interval for advancing non-removable transaction IDs.
5004 : : *
5005 : : * If there is no activity on the node or retention has been stopped, we
5006 : : * progressively double the interval used to advance non-removable transaction
5007 : : * ID. This helps conserve CPU and network resources when there's little benefit
5008 : : * to frequent updates.
5009 : : *
5010 : : * The interval is capped by the lowest of the following:
5011 : : * - wal_receiver_status_interval (if set and retention is active),
5012 : : * - a default maximum of 3 minutes,
5013 : : * - max_retention_duration (if retention is active).
5014 : : *
5015 : : * This ensures the interval never exceeds the retention boundary, even if other
5016 : : * limits are higher. Once activity resumes on the node and the retention is
5017 : : * active, the interval is reset to lesser of 100ms and max_retention_duration,
5018 : : * allowing timely advancement of non-removable transaction ID.
5019 : : *
5020 : : * XXX The use of wal_receiver_status_interval is a bit arbitrary so we can
5021 : : * consider the other interval or a separate GUC if the need arises.
5022 : : */
5023 : : static void
5024 : 148 : adjust_xid_advance_interval(RetainDeadTuplesData *rdt_data, bool new_xid_found)
5025 : : {
5026 [ + + + + ]: 148 : if (rdt_data->xid_advance_interval && !new_xid_found)
5027 : 74 : {
5028 : 74 : int max_interval = wal_receiver_status_interval
5029 : 148 : ? wal_receiver_status_interval * 1000
5030 [ + - ]: 74 : : MAX_XID_ADVANCE_INTERVAL;
5031 : :
5032 : : /*
5033 : : * No new transaction ID has been assigned since the last check, so
5034 : : * double the interval, but not beyond the maximum allowable value.
5035 : : */
5036 : 74 : rdt_data->xid_advance_interval = Min(rdt_data->xid_advance_interval * 2,
5037 : : max_interval);
5038 : : }
5039 [ + + ]: 74 : else if (rdt_data->xid_advance_interval &&
5040 [ + + ]: 59 : !MySubscription->retentionactive)
5041 : : {
5042 : : /*
5043 : : * Retention has been stopped, so double the interval-capped at a
5044 : : * maximum of 3 minutes. The wal_receiver_status_interval is
5045 : : * intentionally not used as an upper bound, since the likelihood of
5046 : : * retention resuming is lower than that of general activity resuming.
5047 : : */
5048 : 1 : rdt_data->xid_advance_interval = Min(rdt_data->xid_advance_interval * 2,
5049 : : MAX_XID_ADVANCE_INTERVAL);
5050 : : }
5051 : : else
5052 : : {
5053 : : /*
5054 : : * A new transaction ID was found or the interval is not yet
5055 : : * initialized, so set the interval to the minimum value.
5056 : : */
5057 : 73 : rdt_data->xid_advance_interval = MIN_XID_ADVANCE_INTERVAL;
5058 : : }
5059 : :
5060 : : /*
5061 : : * Ensure the wait time remains within the maximum retention time limit
5062 : : * when retention is active. Skip this cap when maxretention is zero,
5063 : : * which means unlimited retention (no timeout).
5064 : : */
5065 [ + + - + ]: 148 : if (MySubscription->retentionactive && MySubscription->maxretention > 0)
5066 : 0 : rdt_data->xid_advance_interval = Min(rdt_data->xid_advance_interval,
5067 : : MySubscription->maxretention);
5068 : 148 : }
5069 : :
5070 : : /*
5071 : : * Exit routine for apply workers due to subscription parameter changes.
5072 : : */
5073 : : static void
5074 : 51 : apply_worker_exit(void)
5075 : : {
5076 [ - + ]: 51 : if (am_parallel_apply_worker())
5077 : : {
5078 : : /*
5079 : : * Don't stop the parallel apply worker as the leader will detect the
5080 : : * subscription parameter change and restart logical replication later
5081 : : * anyway. This also prevents the leader from reporting errors when
5082 : : * trying to communicate with a stopped parallel apply worker, which
5083 : : * would accidentally disable subscriptions if disable_on_error was
5084 : : * set.
5085 : : */
5086 : 0 : return;
5087 : : }
5088 : :
5089 : : /*
5090 : : * Reset the last-start time for this apply worker so that the launcher
5091 : : * will restart it without waiting for wal_retrieve_retry_interval if the
5092 : : * subscription is still active, and so that we won't leak that hash table
5093 : : * entry if it isn't.
5094 : : */
5095 [ + - ]: 51 : if (am_leader_apply_worker())
5096 : 51 : ApplyLauncherForgetWorkerStartTime(MyLogicalRepWorker->subid);
5097 : :
5098 : 51 : proc_exit(0);
5099 : : }
5100 : :
5101 : : /*
5102 : : * Reread subscription info if needed.
5103 : : *
5104 : : * For significant changes, we react by exiting the current process; a new
5105 : : * one will be launched afterwards if needed.
5106 : : */
5107 : : void
5108 : 11995 : maybe_reread_subscription(void)
5109 : : {
5110 : : Subscription *newsub;
5111 : : char *old_conninfo;
5112 : : char *new_conninfo;
5113 : 11995 : bool started_tx = false;
5114 : :
5115 : : /* When cache state is valid there is nothing to do here. */
5116 [ + + ]: 11995 : if (MySubscriptionValid)
5117 : 11896 : return;
5118 : :
5119 : : /* This function might be called inside or outside of transaction. */
5120 [ + + ]: 99 : if (!IsTransactionState())
5121 : : {
5122 : 94 : StartTransactionCommand();
5123 : 94 : started_tx = true;
5124 : : }
5125 : :
5126 : 99 : newsub = GetSubscription(MyLogicalRepWorker->subid, true);
5127 : :
5128 [ + - ]: 99 : if (newsub)
5129 : : {
5130 : 99 : MemoryContextSetParent(newsub->cxt, ApplyContext);
5131 : : }
5132 : : else
5133 : : {
5134 : : /*
5135 : : * Exit if the subscription was removed. This normally should not
5136 : : * happen as the worker gets killed during DROP SUBSCRIPTION.
5137 : : */
5138 [ # # ]: 0 : ereport(LOG,
5139 : : (errmsg("logical replication worker for subscription \"%s\" will stop because the subscription was removed",
5140 : : MySubscription->name)));
5141 : :
5142 : : /* Ensure we remove no-longer-useful entry for worker's start time */
5143 [ # # ]: 0 : if (am_leader_apply_worker())
5144 : 0 : ApplyLauncherForgetWorkerStartTime(MyLogicalRepWorker->subid);
5145 : :
5146 : 0 : proc_exit(0);
5147 : : }
5148 : :
5149 : : /* Exit if the subscription was disabled. */
5150 [ + + ]: 99 : if (!newsub->enabled)
5151 : : {
5152 [ + - ]: 15 : ereport(LOG,
5153 : : (errmsg("logical replication worker for subscription \"%s\" will stop because the subscription was disabled",
5154 : : MySubscription->name)));
5155 : :
5156 : 15 : apply_worker_exit();
5157 : : }
5158 : :
5159 : : /*
5160 : : * May raise error, so build conninfo after checking that the subscription
5161 : : * is enabled. Allocated in transaction context; must be copied to
5162 : : * ApplyContext when we set MySubscriptionConninfo.
5163 : : */
5164 : 84 : new_conninfo = SubscriptionConninfo(newsub);
5165 : :
5166 : : /* !slotname should never happen when enabled is true. */
5167 : : Assert(newsub->slotname);
5168 : :
5169 : : /* two-phase cannot be altered while the worker is running */
5170 : : Assert(newsub->twophasestate == MySubscription->twophasestate);
5171 : :
5172 : : /*
5173 : : * Exit if any parameter that affects the remote connection was changed.
5174 : : * The launcher will start a new worker but note that the parallel apply
5175 : : * worker won't restart if the streaming option's value is changed from
5176 : : * 'parallel' to any other value or the server decides not to stream the
5177 : : * in-progress transaction.
5178 : : */
5179 [ + + ]: 84 : if (strcmp(new_conninfo, MySubscriptionConninfo) != 0 ||
5180 [ + + ]: 79 : strcmp(newsub->name, MySubscription->name) != 0 ||
5181 [ + - ]: 78 : strcmp(newsub->slotname, MySubscription->slotname) != 0 ||
5182 [ + + ]: 78 : newsub->binary != MySubscription->binary ||
5183 [ + + ]: 72 : newsub->stream != MySubscription->stream ||
5184 [ + - ]: 67 : newsub->passwordrequired != MySubscription->passwordrequired ||
5185 [ + + ]: 67 : strcmp(newsub->origin, MySubscription->origin) != 0 ||
5186 [ + + ]: 65 : newsub->owner != MySubscription->owner ||
5187 [ + + ]: 64 : !equal(newsub->publications, MySubscription->publications))
5188 : : {
5189 [ - + ]: 30 : if (am_parallel_apply_worker())
5190 [ # # ]: 0 : ereport(LOG,
5191 : : (errmsg("logical replication parallel apply worker for subscription \"%s\" will stop because of a parameter change",
5192 : : MySubscription->name)));
5193 : : else
5194 [ + - ]: 30 : ereport(LOG,
5195 : : (errmsg("logical replication worker for subscription \"%s\" will restart because of a parameter change",
5196 : : MySubscription->name)));
5197 : :
5198 : 30 : apply_worker_exit();
5199 : : }
5200 : :
5201 : : /*
5202 : : * Exit if the subscription owner's superuser privileges have been
5203 : : * revoked.
5204 : : */
5205 [ + + + + ]: 54 : if (!newsub->ownersuperuser && MySubscription->ownersuperuser)
5206 : : {
5207 [ - + ]: 4 : if (am_parallel_apply_worker())
5208 [ # # ]: 0 : ereport(LOG,
5209 : : errmsg("logical replication parallel apply worker for subscription \"%s\" will stop because the subscription owner's superuser privileges have been revoked",
5210 : : MySubscription->name));
5211 : : else
5212 [ + - ]: 4 : ereport(LOG,
5213 : : errmsg("logical replication worker for subscription \"%s\" will restart because the subscription owner's superuser privileges have been revoked",
5214 : : MySubscription->name));
5215 : :
5216 : 4 : apply_worker_exit();
5217 : : }
5218 : :
5219 : : /* Check for other changes that should never happen too. */
5220 [ - + ]: 50 : if (newsub->dbid != MySubscription->dbid)
5221 : : {
5222 [ # # ]: 0 : elog(ERROR, "subscription %u changed unexpectedly",
5223 : : MyLogicalRepWorker->subid);
5224 : : }
5225 : :
5226 : : /* Clean old subscription info and switch to new one. */
5227 : 50 : MemoryContextDelete(MySubscription->cxt);
5228 : 50 : MySubscription = newsub;
5229 : :
5230 : : /* copy to ApplyContext and update MySubscriptionConninfo */
5231 : 50 : old_conninfo = MySubscriptionConninfo;
5232 : 50 : MySubscriptionConninfo = MemoryContextStrdup(ApplyContext, new_conninfo);
5233 : 50 : pfree(old_conninfo);
5234 : :
5235 : : /* Change synchronous commit according to the user's wishes */
5236 : 50 : SetConfigOption("synchronous_commit", MySubscription->synccommit,
5237 : : PGC_BACKEND, PGC_S_OVERRIDE);
5238 : :
5239 : : /* Change wal_receiver_timeout according to the user's wishes */
5240 : 50 : set_wal_receiver_timeout();
5241 : :
5242 [ + + ]: 50 : if (started_tx)
5243 : 48 : CommitTransactionCommand();
5244 : :
5245 : 50 : MySubscriptionValid = true;
5246 : : }
5247 : :
5248 : : /*
5249 : : * Change wal_receiver_timeout to MySubscription->walrcvtimeout.
5250 : : */
5251 : : static void
5252 : 651 : set_wal_receiver_timeout(void)
5253 : : {
5254 : : bool parsed;
5255 : : int val;
5256 : 651 : int prev_timeout = wal_receiver_timeout;
5257 : :
5258 : : /*
5259 : : * Set the wal_receiver_timeout GUC to MySubscription->walrcvtimeout,
5260 : : * which comes from the subscription's wal_receiver_timeout option. If the
5261 : : * value is -1, reset the GUC to its default, meaning it will inherit from
5262 : : * the server config, command line, or role/database settings.
5263 : : */
5264 : 651 : parsed = parse_int(MySubscription->walrcvtimeout, &val, 0, NULL);
5265 [ + - + - ]: 651 : if (parsed && val == -1)
5266 : 651 : SetConfigOption("wal_receiver_timeout", NULL,
5267 : : PGC_BACKEND, PGC_S_SESSION);
5268 : : else
5269 : 0 : SetConfigOption("wal_receiver_timeout", MySubscription->walrcvtimeout,
5270 : : PGC_BACKEND, PGC_S_SESSION);
5271 : :
5272 : : /*
5273 : : * Log the wal_receiver_timeout setting (in milliseconds) as a debug
5274 : : * message when it changes, to verify it was set correctly.
5275 : : */
5276 [ - + ]: 651 : if (prev_timeout != wal_receiver_timeout)
5277 [ # # ]: 0 : elog(DEBUG1, "logical replication worker for subscription \"%s\" wal_receiver_timeout: %d ms",
5278 : : MySubscription->name, wal_receiver_timeout);
5279 : 651 : }
5280 : :
5281 : : /*
5282 : : * Callback from subscription syscache invalidation. Also needed for server or
5283 : : * user mapping invalidation, which can change the connection information for
5284 : : * subscriptions that connect using a server object.
5285 : : */
5286 : : static void
5287 : 103 : subscription_change_cb(Datum arg, SysCacheIdentifier cacheid, uint32 hashvalue)
5288 : : {
5289 : 103 : MySubscriptionValid = false;
5290 : 103 : }
5291 : :
5292 : : /*
5293 : : * subxact_info_write
5294 : : * Store information about subxacts for a toplevel transaction.
5295 : : *
5296 : : * For each subxact we store offset of its first change in the main file.
5297 : : * The file is always over-written as a whole.
5298 : : *
5299 : : * XXX We should only store subxacts that were not aborted yet.
5300 : : */
5301 : : static void
5302 : 371 : subxact_info_write(Oid subid, TransactionId xid)
5303 : : {
5304 : : char path[MAXPGPATH];
5305 : : Size len;
5306 : : BufFile *fd;
5307 : :
5308 : : Assert(TransactionIdIsValid(xid));
5309 : :
5310 : : /* construct the subxact filename */
5311 : 371 : subxact_filename(path, subid, xid);
5312 : :
5313 : : /* Delete the subxacts file, if exists. */
5314 [ + + ]: 371 : if (subxact_data.nsubxacts == 0)
5315 : : {
5316 : 289 : cleanup_subxact_info();
5317 : 289 : BufFileDeleteFileSet(MyLogicalRepWorker->stream_fileset, path, true);
5318 : :
5319 : 289 : return;
5320 : : }
5321 : :
5322 : : /*
5323 : : * Create the subxact file if it not already created, otherwise open the
5324 : : * existing file.
5325 : : */
5326 : 82 : fd = BufFileOpenFileSet(MyLogicalRepWorker->stream_fileset, path, O_RDWR,
5327 : : true);
5328 [ + + ]: 82 : if (fd == NULL)
5329 : 8 : fd = BufFileCreateFileSet(MyLogicalRepWorker->stream_fileset, path);
5330 : :
5331 : 82 : len = sizeof(SubXactInfo) * subxact_data.nsubxacts;
5332 : :
5333 : : /* Write the subxact count and subxact info */
5334 : 82 : BufFileWrite(fd, &subxact_data.nsubxacts, sizeof(subxact_data.nsubxacts));
5335 : 82 : BufFileWrite(fd, subxact_data.subxacts, len);
5336 : :
5337 : 82 : BufFileClose(fd);
5338 : :
5339 : : /* free the memory allocated for subxact info */
5340 : 82 : cleanup_subxact_info();
5341 : : }
5342 : :
5343 : : /*
5344 : : * subxact_info_read
5345 : : * Restore information about subxacts of a streamed transaction.
5346 : : *
5347 : : * Read information about subxacts into the structure subxact_data that can be
5348 : : * used later.
5349 : : */
5350 : : static void
5351 : 343 : subxact_info_read(Oid subid, TransactionId xid)
5352 : : {
5353 : : char path[MAXPGPATH];
5354 : : Size len;
5355 : : BufFile *fd;
5356 : : MemoryContext oldctx;
5357 : :
5358 : : Assert(!subxact_data.subxacts);
5359 : : Assert(subxact_data.nsubxacts == 0);
5360 : : Assert(subxact_data.nsubxacts_max == 0);
5361 : :
5362 : : /*
5363 : : * If the subxact file doesn't exist that means we don't have any subxact
5364 : : * info.
5365 : : */
5366 : 343 : subxact_filename(path, subid, xid);
5367 : 343 : fd = BufFileOpenFileSet(MyLogicalRepWorker->stream_fileset, path, O_RDONLY,
5368 : : true);
5369 [ + + ]: 343 : if (fd == NULL)
5370 : 264 : return;
5371 : :
5372 : : /* read number of subxact items */
5373 : 79 : BufFileReadExact(fd, &subxact_data.nsubxacts, sizeof(subxact_data.nsubxacts));
5374 : :
5375 : 79 : len = sizeof(SubXactInfo) * subxact_data.nsubxacts;
5376 : :
5377 : : /* we keep the maximum as a power of 2 */
5378 : 79 : subxact_data.nsubxacts_max = 1 << pg_ceil_log2_32(subxact_data.nsubxacts);
5379 : :
5380 : : /*
5381 : : * Allocate subxact information in the logical streaming context. We need
5382 : : * this information during the complete stream so that we can add the sub
5383 : : * transaction info to this. On stream stop we will flush this information
5384 : : * to the subxact file and reset the logical streaming context.
5385 : : */
5386 : 79 : oldctx = MemoryContextSwitchTo(LogicalStreamingContext);
5387 : 79 : subxact_data.subxacts = palloc_array(SubXactInfo,
5388 : : subxact_data.nsubxacts_max);
5389 : 79 : MemoryContextSwitchTo(oldctx);
5390 : :
5391 [ + - ]: 79 : if (len > 0)
5392 : 79 : BufFileReadExact(fd, subxact_data.subxacts, len);
5393 : :
5394 : 79 : BufFileClose(fd);
5395 : : }
5396 : :
5397 : : /*
5398 : : * subxact_info_add
5399 : : * Add information about a subxact (offset in the main file).
5400 : : */
5401 : : static void
5402 : 102513 : subxact_info_add(TransactionId xid)
5403 : : {
5404 : 102513 : SubXactInfo *subxacts = subxact_data.subxacts;
5405 : : int64 i;
5406 : :
5407 : : /* We must have a valid top level stream xid and a stream fd. */
5408 : : Assert(TransactionIdIsValid(stream_xid));
5409 : : Assert(stream_fd != NULL);
5410 : :
5411 : : /*
5412 : : * If the XID matches the toplevel transaction, we don't want to add it.
5413 : : */
5414 [ + + ]: 102513 : if (stream_xid == xid)
5415 : 92389 : return;
5416 : :
5417 : : /*
5418 : : * In most cases we're checking the same subxact as we've already seen in
5419 : : * the last call, so make sure to ignore it (this change comes later).
5420 : : */
5421 [ + + ]: 10124 : if (subxact_data.subxact_last == xid)
5422 : 10048 : return;
5423 : :
5424 : : /* OK, remember we're processing this XID. */
5425 : 76 : subxact_data.subxact_last = xid;
5426 : :
5427 : : /*
5428 : : * Check if the transaction is already present in the array of subxact. We
5429 : : * intentionally scan the array from the tail, because we're likely adding
5430 : : * a change for the most recent subtransactions.
5431 : : *
5432 : : * XXX Can we rely on the subxact XIDs arriving in sorted order? That
5433 : : * would allow us to use binary search here.
5434 : : */
5435 [ + + ]: 95 : for (i = subxact_data.nsubxacts; i > 0; i--)
5436 : : {
5437 : : /* found, so we're done */
5438 [ + + ]: 76 : if (subxacts[i - 1].xid == xid)
5439 : 57 : return;
5440 : : }
5441 : :
5442 : : /* This is a new subxact, so we need to add it to the array. */
5443 [ + + ]: 19 : if (subxact_data.nsubxacts == 0)
5444 : : {
5445 : : MemoryContext oldctx;
5446 : :
5447 : 8 : subxact_data.nsubxacts_max = 128;
5448 : :
5449 : : /*
5450 : : * Allocate this memory for subxacts in per-stream context, see
5451 : : * subxact_info_read.
5452 : : */
5453 : 8 : oldctx = MemoryContextSwitchTo(LogicalStreamingContext);
5454 : 8 : subxacts = palloc_array(SubXactInfo, subxact_data.nsubxacts_max);
5455 : 8 : MemoryContextSwitchTo(oldctx);
5456 : : }
5457 [ + + ]: 11 : else if (subxact_data.nsubxacts == subxact_data.nsubxacts_max)
5458 : : {
5459 : 10 : subxact_data.nsubxacts_max *= 2;
5460 : 10 : subxacts = repalloc_array(subxacts, SubXactInfo,
5461 : : subxact_data.nsubxacts_max);
5462 : : }
5463 : :
5464 : 19 : subxacts[subxact_data.nsubxacts].xid = xid;
5465 : :
5466 : : /*
5467 : : * Get the current offset of the stream file and store it as offset of
5468 : : * this subxact.
5469 : : */
5470 : 19 : BufFileTell(stream_fd,
5471 : 19 : &subxacts[subxact_data.nsubxacts].fileno,
5472 : 19 : &subxacts[subxact_data.nsubxacts].offset);
5473 : :
5474 : 19 : subxact_data.nsubxacts++;
5475 : 19 : subxact_data.subxacts = subxacts;
5476 : : }
5477 : :
5478 : : /* format filename for file containing the info about subxacts */
5479 : : static inline void
5480 : 745 : subxact_filename(char *path, Oid subid, TransactionId xid)
5481 : : {
5482 : 745 : snprintf(path, MAXPGPATH, "%u-%u.subxacts", subid, xid);
5483 : 745 : }
5484 : :
5485 : : /* format filename for file containing serialized changes */
5486 : : static inline void
5487 : 437 : changes_filename(char *path, Oid subid, TransactionId xid)
5488 : : {
5489 : 437 : snprintf(path, MAXPGPATH, "%u-%u.changes", subid, xid);
5490 : 437 : }
5491 : :
5492 : : /*
5493 : : * stream_cleanup_files
5494 : : * Cleanup files for a subscription / toplevel transaction.
5495 : : *
5496 : : * Remove files with serialized changes and subxact info for a particular
5497 : : * toplevel transaction. Each subscription has a separate set of files
5498 : : * for any toplevel transaction.
5499 : : */
5500 : : void
5501 : 31 : stream_cleanup_files(Oid subid, TransactionId xid)
5502 : : {
5503 : : char path[MAXPGPATH];
5504 : :
5505 : : /* Delete the changes file. */
5506 : 31 : changes_filename(path, subid, xid);
5507 : 31 : BufFileDeleteFileSet(MyLogicalRepWorker->stream_fileset, path, false);
5508 : :
5509 : : /* Delete the subxact file, if it exists. */
5510 : 31 : subxact_filename(path, subid, xid);
5511 : 31 : BufFileDeleteFileSet(MyLogicalRepWorker->stream_fileset, path, true);
5512 : 31 : }
5513 : :
5514 : : /*
5515 : : * stream_open_file
5516 : : * Open a file that we'll use to serialize changes for a toplevel
5517 : : * transaction.
5518 : : *
5519 : : * Open a file for streamed changes from a toplevel transaction identified
5520 : : * by stream_xid (global variable). If it's the first chunk of streamed
5521 : : * changes for this transaction, create the buffile, otherwise open the
5522 : : * previously created file.
5523 : : */
5524 : : static void
5525 : 362 : stream_open_file(Oid subid, TransactionId xid, bool first_segment)
5526 : : {
5527 : : char path[MAXPGPATH];
5528 : : MemoryContext oldcxt;
5529 : :
5530 : : Assert(OidIsValid(subid));
5531 : : Assert(TransactionIdIsValid(xid));
5532 : : Assert(stream_fd == NULL);
5533 : :
5534 : :
5535 : 362 : changes_filename(path, subid, xid);
5536 [ - + ]: 362 : elog(DEBUG1, "opening file \"%s\" for streamed changes", path);
5537 : :
5538 : : /*
5539 : : * Create/open the buffiles under the logical streaming context so that we
5540 : : * have those files until stream stop.
5541 : : */
5542 : 362 : oldcxt = MemoryContextSwitchTo(LogicalStreamingContext);
5543 : :
5544 : : /*
5545 : : * If this is the first streamed segment, create the changes file.
5546 : : * Otherwise, just open the file for writing, in append mode.
5547 : : */
5548 [ + + ]: 362 : if (first_segment)
5549 : 32 : stream_fd = BufFileCreateFileSet(MyLogicalRepWorker->stream_fileset,
5550 : : path);
5551 : : else
5552 : : {
5553 : : /*
5554 : : * Open the file and seek to the end of the file because we always
5555 : : * append the changes file.
5556 : : */
5557 : 330 : stream_fd = BufFileOpenFileSet(MyLogicalRepWorker->stream_fileset,
5558 : : path, O_RDWR, false);
5559 : 330 : BufFileSeek(stream_fd, 0, 0, SEEK_END);
5560 : : }
5561 : :
5562 : 362 : MemoryContextSwitchTo(oldcxt);
5563 : 362 : }
5564 : :
5565 : : /*
5566 : : * stream_close_file
5567 : : * Close the currently open file with streamed changes.
5568 : : */
5569 : : static void
5570 : 392 : stream_close_file(void)
5571 : : {
5572 : : Assert(stream_fd != NULL);
5573 : :
5574 : 392 : BufFileClose(stream_fd);
5575 : :
5576 : 392 : stream_fd = NULL;
5577 : 392 : }
5578 : :
5579 : : /*
5580 : : * stream_write_change
5581 : : * Serialize a change to a file for the current toplevel transaction.
5582 : : *
5583 : : * The change is serialized in a simple format, with length (not including
5584 : : * the length), action code (identifying the message type) and message
5585 : : * contents (without the subxact TransactionId value).
5586 : : */
5587 : : static void
5588 : 107554 : stream_write_change(char action, StringInfo s)
5589 : : {
5590 : : int len;
5591 : :
5592 : : Assert(stream_fd != NULL);
5593 : :
5594 : : /* total on-disk size, including the action type character */
5595 : 107554 : len = (s->len - s->cursor) + sizeof(char);
5596 : :
5597 : : /* first write the size */
5598 : 107554 : BufFileWrite(stream_fd, &len, sizeof(len));
5599 : :
5600 : : /* then the action */
5601 : 107554 : BufFileWrite(stream_fd, &action, sizeof(action));
5602 : :
5603 : : /* and finally the remaining part of the buffer (after the XID) */
5604 : 107554 : len = (s->len - s->cursor);
5605 : :
5606 : 107554 : BufFileWrite(stream_fd, &s->data[s->cursor], len);
5607 : 107554 : }
5608 : :
5609 : : /*
5610 : : * stream_open_and_write_change
5611 : : * Serialize a message to a file for the given transaction.
5612 : : *
5613 : : * This function is similar to stream_write_change except that it will open the
5614 : : * target file if not already before writing the message and close the file at
5615 : : * the end.
5616 : : */
5617 : : static void
5618 : 5 : stream_open_and_write_change(TransactionId xid, char action, StringInfo s)
5619 : : {
5620 : : Assert(!in_streamed_transaction);
5621 : :
5622 [ + - ]: 5 : if (!stream_fd)
5623 : 5 : stream_start_internal(xid, false);
5624 : :
5625 : 5 : stream_write_change(action, s);
5626 : 5 : stream_stop_internal(xid);
5627 : 5 : }
5628 : :
5629 : : /*
5630 : : * Sets streaming options including replication slot name and origin start
5631 : : * position. Workers need these options for logical replication.
5632 : : */
5633 : : void
5634 : 496 : set_stream_options(WalRcvStreamOptions *options,
5635 : : char *slotname,
5636 : : XLogRecPtr *origin_startpos)
5637 : : {
5638 : : int server_version;
5639 : :
5640 : 496 : options->logical = true;
5641 : 496 : options->startpoint = *origin_startpos;
5642 : 496 : options->slotname = slotname;
5643 : :
5644 : 496 : server_version = walrcv_server_version(LogRepWorkerWalRcvConn);
5645 : 496 : options->proto.logical.proto_version =
5646 [ - + - - : 496 : server_version >= 160000 ? LOGICALREP_PROTO_STREAM_PARALLEL_VERSION_NUM :
- - ]
5647 : : server_version >= 150000 ? LOGICALREP_PROTO_TWOPHASE_VERSION_NUM :
5648 : : server_version >= 140000 ? LOGICALREP_PROTO_STREAM_VERSION_NUM :
5649 : : LOGICALREP_PROTO_VERSION_NUM;
5650 : :
5651 : 496 : options->proto.logical.publication_names = MySubscription->publications;
5652 : 496 : options->proto.logical.binary = MySubscription->binary;
5653 : :
5654 : : /*
5655 : : * Assign the appropriate option value for streaming option according to
5656 : : * the 'streaming' mode and the publisher's ability to support that mode.
5657 : : */
5658 [ + - ]: 496 : if (server_version >= 160000 &&
5659 [ + + ]: 496 : MySubscription->stream == LOGICALREP_STREAM_PARALLEL)
5660 : : {
5661 : 462 : options->proto.logical.streaming_str = "parallel";
5662 : 462 : MyLogicalRepWorker->parallel_apply = true;
5663 : : }
5664 [ + - ]: 34 : else if (server_version >= 140000 &&
5665 [ + + ]: 34 : MySubscription->stream != LOGICALREP_STREAM_OFF)
5666 : : {
5667 : 26 : options->proto.logical.streaming_str = "on";
5668 : 26 : MyLogicalRepWorker->parallel_apply = false;
5669 : : }
5670 : : else
5671 : : {
5672 : 8 : options->proto.logical.streaming_str = NULL;
5673 : 8 : MyLogicalRepWorker->parallel_apply = false;
5674 : : }
5675 : :
5676 : 496 : options->proto.logical.twophase = false;
5677 : 496 : options->proto.logical.origin = pstrdup(MySubscription->origin);
5678 : 496 : }
5679 : :
5680 : : /*
5681 : : * Cleanup the memory for subxacts and reset the related variables.
5682 : : */
5683 : : static inline void
5684 : 375 : cleanup_subxact_info(void)
5685 : : {
5686 [ + + ]: 375 : if (subxact_data.subxacts)
5687 : 87 : pfree(subxact_data.subxacts);
5688 : :
5689 : 375 : subxact_data.subxacts = NULL;
5690 : 375 : subxact_data.subxact_last = InvalidTransactionId;
5691 : 375 : subxact_data.nsubxacts = 0;
5692 : 375 : subxact_data.nsubxacts_max = 0;
5693 : 375 : }
5694 : :
5695 : : /*
5696 : : * Common function to run the apply loop with error handling. Disable the
5697 : : * subscription, if necessary.
5698 : : *
5699 : : * Note that we don't handle FATAL errors which are probably because
5700 : : * of system resource error and are not repeatable.
5701 : : */
5702 : : void
5703 : 495 : start_apply(XLogRecPtr origin_startpos)
5704 : : {
5705 [ + + ]: 495 : PG_TRY();
5706 : : {
5707 : 495 : LogicalRepApplyLoop(origin_startpos);
5708 : : }
5709 : 128 : PG_CATCH();
5710 : : {
5711 : : /*
5712 : : * Reset the origin state to prevent the advancement of origin
5713 : : * progress if we fail to apply. Otherwise, this will result in
5714 : : * transaction loss as that transaction won't be sent again by the
5715 : : * server.
5716 : : */
5717 : 128 : replorigin_xact_clear(true);
5718 : :
5719 [ + + ]: 128 : if (MySubscription->disableonerr)
5720 : 3 : DisableSubscriptionAndExit();
5721 : : else
5722 : : {
5723 : : /*
5724 : : * Report the worker failed while applying changes. Abort the
5725 : : * current transaction so that the stats message is sent in an
5726 : : * idle state.
5727 : : */
5728 : 125 : AbortOutOfAnyTransaction();
5729 : 125 : pgstat_report_subscription_error(MySubscription->oid);
5730 : :
5731 : 125 : PG_RE_THROW();
5732 : : }
5733 : : }
5734 [ # # ]: 0 : PG_END_TRY();
5735 : 0 : }
5736 : :
5737 : : /*
5738 : : * Runs the leader apply worker.
5739 : : *
5740 : : * It sets up replication origin, streaming options and then starts streaming.
5741 : : */
5742 : : static void
5743 : 352 : run_apply_worker(void)
5744 : : {
5745 : : char originname[NAMEDATALEN];
5746 : 352 : XLogRecPtr origin_startpos = InvalidXLogRecPtr;
5747 : 352 : char *slotname = NULL;
5748 : : WalRcvStreamOptions options;
5749 : : ReplOriginId originid;
5750 : : TimeLineID startpointTLI;
5751 : : char *err;
5752 : : bool must_use_password;
5753 : :
5754 : 352 : slotname = MySubscription->slotname;
5755 : :
5756 : : /*
5757 : : * This shouldn't happen if the subscription is enabled, but guard against
5758 : : * DDL bugs or manual catalog changes. (libpqwalreceiver will crash if
5759 : : * slot is NULL.)
5760 : : */
5761 [ - + ]: 352 : if (!slotname)
5762 [ # # ]: 0 : ereport(ERROR,
5763 : : (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
5764 : : errmsg("subscription has no replication slot set")));
5765 : :
5766 : : /* Setup replication origin tracking. */
5767 : 352 : ReplicationOriginNameForLogicalRep(MySubscription->oid, InvalidOid,
5768 : : originname, sizeof(originname));
5769 : 352 : StartTransactionCommand();
5770 : 352 : originid = replorigin_by_name(originname, true);
5771 [ - + ]: 352 : if (!OidIsValid(originid))
5772 : 0 : originid = replorigin_create(originname);
5773 : 352 : replorigin_session_setup(originid, 0);
5774 : 352 : replorigin_xact_state.origin = originid;
5775 : 352 : origin_startpos = replorigin_session_get_progress(false);
5776 : 352 : CommitTransactionCommand();
5777 : :
5778 : : /* Is the use of a password mandatory? */
5779 [ + + ]: 675 : must_use_password = MySubscription->passwordrequired &&
5780 [ + + ]: 323 : !MySubscription->ownersuperuser;
5781 : :
5782 : 352 : LogRepWorkerWalRcvConn = walrcv_connect(MySubscriptionConninfo, true,
5783 : : true, must_use_password,
5784 : : MySubscription->name, &err);
5785 : :
5786 [ + + ]: 335 : if (LogRepWorkerWalRcvConn == NULL)
5787 [ + - ]: 44 : ereport(ERROR,
5788 : : (errcode(ERRCODE_CONNECTION_FAILURE),
5789 : : errmsg("apply worker for subscription \"%s\" could not connect to the publisher: %s",
5790 : : MySubscription->name, err)));
5791 : :
5792 : : /*
5793 : : * We don't really use the output identify_system for anything but it does
5794 : : * some initializations on the upstream so let's still call it.
5795 : : */
5796 : 291 : (void) walrcv_identify_system(LogRepWorkerWalRcvConn, &startpointTLI, NULL);
5797 : :
5798 : : /*
5799 : : * If retain_dead_tuples is enabled, verify that the publisher is
5800 : : * suitable, that is, it runs a version that supports the feature and is
5801 : : * not in recovery. This is the authoritative check. Although the same
5802 : : * validation is performed opportunistically at DDL time, the publisher's
5803 : : * version or recovery status may have changed since then, for example
5804 : : * after a failover.
5805 : : */
5806 [ + + ]: 291 : if (MySubscription->retaindeadtuples)
5807 : : {
5808 : 15 : StartTransactionCommand();
5809 : 15 : CheckPubDeadTupleRetention(LogRepWorkerWalRcvConn);
5810 : 15 : CommitTransactionCommand();
5811 : : }
5812 : :
5813 : 291 : set_apply_error_context_origin(originname);
5814 : :
5815 : 291 : set_stream_options(&options, slotname, &origin_startpos);
5816 : :
5817 : : /*
5818 : : * Even when the two_phase mode is requested by the user, it remains as
5819 : : * the tri-state PENDING until all tablesyncs have reached READY state.
5820 : : * Only then, can it become ENABLED.
5821 : : *
5822 : : * Note: If the subscription has no tables then leave the state as
5823 : : * PENDING, which allows ALTER SUBSCRIPTION ... REFRESH PUBLICATION to
5824 : : * work.
5825 : : */
5826 [ + + + + ]: 307 : if (MySubscription->twophasestate == LOGICALREP_TWOPHASE_STATE_PENDING &&
5827 : 16 : AllTablesyncsReady())
5828 : : {
5829 : : /* Start streaming with two_phase enabled */
5830 : 9 : options.proto.logical.twophase = true;
5831 : 9 : walrcv_startstreaming(LogRepWorkerWalRcvConn, &options);
5832 : :
5833 : 9 : StartTransactionCommand();
5834 : :
5835 : : /*
5836 : : * Updating pg_subscription might involve TOAST table access, so
5837 : : * ensure we have a valid snapshot.
5838 : : */
5839 : 9 : PushActiveSnapshot(GetTransactionSnapshot());
5840 : :
5841 : 9 : UpdateTwoPhaseState(MySubscription->oid, LOGICALREP_TWOPHASE_STATE_ENABLED);
5842 : 9 : MySubscription->twophasestate = LOGICALREP_TWOPHASE_STATE_ENABLED;
5843 : 9 : PopActiveSnapshot();
5844 : 9 : CommitTransactionCommand();
5845 : : }
5846 : : else
5847 : : {
5848 : 282 : walrcv_startstreaming(LogRepWorkerWalRcvConn, &options);
5849 : : }
5850 : :
5851 [ + + + + : 290 : ereport(DEBUG1,
+ - + - ]
5852 : : (errmsg_internal("logical replication apply worker for subscription \"%s\" two_phase is %s",
5853 : : MySubscription->name,
5854 : : MySubscription->twophasestate == LOGICALREP_TWOPHASE_STATE_DISABLED ? "DISABLED" :
5855 : : MySubscription->twophasestate == LOGICALREP_TWOPHASE_STATE_PENDING ? "PENDING" :
5856 : : MySubscription->twophasestate == LOGICALREP_TWOPHASE_STATE_ENABLED ? "ENABLED" :
5857 : : "?")));
5858 : :
5859 : : /* Run the main loop. */
5860 : 290 : start_apply(origin_startpos);
5861 : 0 : }
5862 : :
5863 : : /*
5864 : : * Common initialization for leader apply worker, parallel apply worker,
5865 : : * tablesync worker and sequencesync worker.
5866 : : *
5867 : : * Initialize the database connection, in-memory subscription and necessary
5868 : : * config options.
5869 : : */
5870 : : void
5871 : 680 : InitializeLogRepWorker(void)
5872 : : {
5873 : : /* Run as replica session replication role. */
5874 : 680 : SetConfigOption("session_replication_role", "replica",
5875 : : PGC_SUSET, PGC_S_OVERRIDE);
5876 : :
5877 : : /* Connect to our database. */
5878 : 680 : BackgroundWorkerInitializeConnectionByOid(MyLogicalRepWorker->dbid,
5879 : 680 : MyLogicalRepWorker->userid,
5880 : : 0);
5881 : :
5882 : : /*
5883 : : * Set always-secure search path, so malicious users can't redirect user
5884 : : * code (e.g. pg_index.indexprs).
5885 : : */
5886 : 674 : SetConfigOption("search_path", "", PGC_SUSET, PGC_S_OVERRIDE);
5887 : :
5888 : : /*
5889 : : * Ignore default_transaction_read_only for logical replication workers,
5890 : : * as they need to be able to modify subscriber-side state regardless of
5891 : : * that setting.
5892 : : */
5893 : 674 : SetConfigOption("default_transaction_read_only", "off", PGC_SUSET,
5894 : : PGC_S_OVERRIDE);
5895 : :
5896 : 674 : ApplyContext = AllocSetContextCreate(TopMemoryContext,
5897 : : "ApplyContext",
5898 : : ALLOCSET_DEFAULT_SIZES);
5899 : :
5900 : 674 : StartTransactionCommand();
5901 : :
5902 : : /*
5903 : : * Lock the subscription to prevent it from being concurrently dropped,
5904 : : * then re-verify its existence. After the initialization, the worker will
5905 : : * be terminated gracefully if the subscription is dropped.
5906 : : */
5907 : 674 : LockSharedObject(SubscriptionRelationId, MyLogicalRepWorker->subid, 0,
5908 : : AccessShareLock);
5909 : :
5910 : 673 : MySubscription = GetSubscription(MyLogicalRepWorker->subid, true);
5911 : :
5912 [ + + ]: 673 : if (MySubscription)
5913 : : {
5914 : 602 : MemoryContextSetParent(MySubscription->cxt, ApplyContext);
5915 : : }
5916 : : else
5917 : : {
5918 [ + - ]: 71 : ereport(LOG,
5919 : : (errmsg("logical replication worker for subscription %u will not start because the subscription was removed during startup",
5920 : : MyLogicalRepWorker->subid)));
5921 : :
5922 : : /* Ensure we remove no-longer-useful entry for worker's start time */
5923 [ + - ]: 71 : if (am_leader_apply_worker())
5924 : 71 : ApplyLauncherForgetWorkerStartTime(MyLogicalRepWorker->subid);
5925 : :
5926 : 71 : proc_exit(0);
5927 : : }
5928 : :
5929 [ + + ]: 602 : if (!MySubscription->enabled)
5930 : : {
5931 [ + - ]: 1 : ereport(LOG,
5932 : : (errmsg("logical replication worker for subscription \"%s\" will not start because the subscription was disabled during startup",
5933 : : MySubscription->name)));
5934 : :
5935 : 1 : apply_worker_exit();
5936 : : }
5937 : :
5938 : : /*
5939 : : * May raise error for server-based subscriptions, so build conninfo after
5940 : : * checking that the subscription is enabled. Build in transaction context
5941 : : * and copy to ApplyContext.
5942 : : */
5943 : 601 : MySubscriptionConninfo =
5944 : 601 : MemoryContextStrdup(ApplyContext,
5945 : 601 : SubscriptionConninfo(MySubscription));
5946 : :
5947 : 601 : MySubscriptionValid = true;
5948 : :
5949 : : /*
5950 : : * Restart the worker if retain_dead_tuples was enabled during startup.
5951 : : *
5952 : : * At this point, the replication slot used for conflict detection might
5953 : : * not exist yet, or could be dropped soon if the launcher perceives
5954 : : * retain_dead_tuples as disabled. To avoid unnecessary tracking of
5955 : : * oldest_nonremovable_xid when the slot is absent or at risk of being
5956 : : * dropped, a restart is initiated.
5957 : : *
5958 : : * The oldest_nonremovable_xid should be initialized only when the
5959 : : * subscription's retention is active before launching the worker. See
5960 : : * logicalrep_worker_launch.
5961 : : */
5962 [ + + ]: 601 : if (am_leader_apply_worker() &&
5963 [ + + ]: 352 : MySubscription->retaindeadtuples &&
5964 [ + - ]: 17 : MySubscription->retentionactive &&
5965 [ - + ]: 17 : !TransactionIdIsValid(MyLogicalRepWorker->oldest_nonremovable_xid))
5966 : : {
5967 [ # # ]: 0 : ereport(LOG,
5968 : : errmsg("logical replication worker for subscription \"%s\" will restart because the option %s was enabled during startup",
5969 : : MySubscription->name, "retain_dead_tuples"));
5970 : :
5971 : 0 : apply_worker_exit();
5972 : : }
5973 : :
5974 : : /* Setup synchronous commit according to the user's wishes */
5975 : 601 : SetConfigOption("synchronous_commit", MySubscription->synccommit,
5976 : : PGC_BACKEND, PGC_S_OVERRIDE);
5977 : :
5978 : : /* Change wal_receiver_timeout according to the user's wishes */
5979 : 601 : set_wal_receiver_timeout();
5980 : :
5981 : : /*
5982 : : * Keep us informed about subscription or role changes. Note that the
5983 : : * role's superuser privilege can be revoked.
5984 : : */
5985 : 601 : CacheRegisterSyscacheCallback(SUBSCRIPTIONOID,
5986 : : subscription_change_cb,
5987 : : (Datum) 0);
5988 : : /* Changes to foreign servers may affect subscriptions using SERVER. */
5989 : 601 : CacheRegisterSyscacheCallback(FOREIGNSERVEROID,
5990 : : subscription_change_cb,
5991 : : (Datum) 0);
5992 : : /* Changes to user mappings may affect subscriptions using SERVER. */
5993 : 601 : CacheRegisterSyscacheCallback(USERMAPPINGOID,
5994 : : subscription_change_cb,
5995 : : (Datum) 0);
5996 : :
5997 : : /*
5998 : : * Changes to FDW connection_function may affect subscriptions using
5999 : : * SERVER.
6000 : : */
6001 : 601 : CacheRegisterSyscacheCallback(FOREIGNDATAWRAPPEROID,
6002 : : subscription_change_cb,
6003 : : (Datum) 0);
6004 : :
6005 : 601 : CacheRegisterSyscacheCallback(AUTHOID,
6006 : : subscription_change_cb,
6007 : : (Datum) 0);
6008 : :
6009 [ + + ]: 601 : if (am_tablesync_worker())
6010 [ + - ]: 222 : ereport(LOG,
6011 : : errmsg("logical replication table synchronization worker for subscription \"%s\", table \"%s\" has started",
6012 : : MySubscription->name,
6013 : : get_rel_name(MyLogicalRepWorker->relid)));
6014 [ + + ]: 379 : else if (am_sequencesync_worker())
6015 [ + - ]: 15 : ereport(LOG,
6016 : : errmsg("logical replication sequence synchronization worker for subscription \"%s\" has started",
6017 : : MySubscription->name));
6018 : : else
6019 [ + - ]: 364 : ereport(LOG,
6020 : : errmsg("logical replication apply worker for subscription \"%s\" has started",
6021 : : MySubscription->name));
6022 : :
6023 : 601 : CommitTransactionCommand();
6024 : :
6025 : : /*
6026 : : * Register a callback to reset the origin state before aborting any
6027 : : * pending transaction during shutdown (see ShutdownPostgres()). This will
6028 : : * avoid origin advancement for an incomplete transaction which could
6029 : : * otherwise lead to its loss as such a transaction won't be sent by the
6030 : : * server again.
6031 : : *
6032 : : * Note that even a LOG or DEBUG statement placed after setting the origin
6033 : : * state may process a shutdown signal before committing the current apply
6034 : : * operation. So, it is important to register such a callback here.
6035 : : *
6036 : : * Register this callback here to ensure that all types of logical
6037 : : * replication workers that set up origins and apply remote transactions
6038 : : * are protected.
6039 : : */
6040 : 601 : before_shmem_exit(on_exit_clear_xact_state, (Datum) 0);
6041 : 601 : }
6042 : :
6043 : : /*
6044 : : * Callback on exit to clear transaction-level replication origin state.
6045 : : */
6046 : : static void
6047 : 601 : on_exit_clear_xact_state(int code, Datum arg)
6048 : : {
6049 : 601 : replorigin_xact_clear(true);
6050 : 601 : }
6051 : :
6052 : : /*
6053 : : * Common function to setup the leader apply, tablesync and sequencesync worker.
6054 : : */
6055 : : void
6056 : 668 : SetupApplyOrSyncWorker(int worker_slot)
6057 : : {
6058 : : /* Attach to slot */
6059 : 668 : logicalrep_worker_attach(worker_slot);
6060 : :
6061 : : Assert(am_tablesync_worker() || am_sequencesync_worker() || am_leader_apply_worker());
6062 : :
6063 : : /* Setup signal handling */
6064 : 668 : pqsignal(SIGHUP, SignalHandlerForConfigReload);
6065 : 668 : BackgroundWorkerUnblockSignals();
6066 : :
6067 : : /*
6068 : : * We don't currently need any ResourceOwner in a walreceiver process, but
6069 : : * if we did, we could call CreateAuxProcessResourceOwner here.
6070 : : */
6071 : :
6072 : : /* Initialise stats to a sanish value */
6073 [ + + ]: 668 : if (am_sequencesync_worker())
6074 : : {
6075 : 15 : MyLogicalRepWorker->last_send_time =
6076 : 15 : MyLogicalRepWorker->last_recv_time =
6077 : 15 : MyLogicalRepWorker->reply_time = 0;
6078 : : }
6079 : : else
6080 : : {
6081 : 653 : MyLogicalRepWorker->last_send_time =
6082 : 653 : MyLogicalRepWorker->last_recv_time =
6083 : 653 : MyLogicalRepWorker->reply_time = GetCurrentTimestamp();
6084 : : }
6085 : :
6086 : : /* Load the libpq-specific functions */
6087 : 668 : load_file("libpqwalreceiver", false);
6088 : :
6089 : 668 : InitializeLogRepWorker();
6090 : :
6091 : : /*
6092 : : * Setup callback for syscache so that we know when something changes in
6093 : : * the subscription relation state.
6094 : : */
6095 : 589 : CacheRegisterSyscacheCallback(SUBSCRIPTIONRELMAP,
6096 : : InvalidateSyncingRelStates,
6097 : : (Datum) 0);
6098 : 589 : }
6099 : :
6100 : : /* Logical Replication Apply worker entry point */
6101 : : void
6102 : 431 : ApplyWorkerMain(Datum main_arg)
6103 : : {
6104 : 431 : int worker_slot = DatumGetInt32(main_arg);
6105 : :
6106 : 431 : InitializingApplyWorker = true;
6107 : :
6108 : 431 : SetupApplyOrSyncWorker(worker_slot);
6109 : :
6110 : 352 : InitializingApplyWorker = false;
6111 : :
6112 : 352 : run_apply_worker();
6113 : :
6114 : 0 : proc_exit(0);
6115 : : }
6116 : :
6117 : : /*
6118 : : * After error recovery, disable the subscription in a new transaction
6119 : : * and exit cleanly.
6120 : : */
6121 : : void
6122 : 4 : DisableSubscriptionAndExit(void)
6123 : : {
6124 : : /*
6125 : : * Emit the error message, and recover from the error state to an idle
6126 : : * state
6127 : : */
6128 : 4 : HOLD_INTERRUPTS();
6129 : :
6130 : 4 : EmitErrorReport();
6131 : 4 : AbortOutOfAnyTransaction();
6132 : 4 : FlushErrorState();
6133 : :
6134 : 4 : RESUME_INTERRUPTS();
6135 : :
6136 : : /*
6137 : : * Report the worker failed during sequence synchronization, table
6138 : : * synchronization, or apply.
6139 : : */
6140 : 4 : pgstat_report_subscription_error(MyLogicalRepWorker->subid);
6141 : :
6142 : : /* Disable the subscription */
6143 : 4 : StartTransactionCommand();
6144 : :
6145 : : /*
6146 : : * Updating pg_subscription might involve TOAST table access, so ensure we
6147 : : * have a valid snapshot.
6148 : : */
6149 : 4 : PushActiveSnapshot(GetTransactionSnapshot());
6150 : :
6151 : 4 : DisableSubscription(MySubscription->oid);
6152 : 4 : PopActiveSnapshot();
6153 : 4 : CommitTransactionCommand();
6154 : :
6155 : : /* Ensure we remove no-longer-useful entry for worker's start time */
6156 [ + + ]: 4 : if (am_leader_apply_worker())
6157 : 3 : ApplyLauncherForgetWorkerStartTime(MyLogicalRepWorker->subid);
6158 : :
6159 : : /* Notify the subscription has been disabled and exit */
6160 [ + - ]: 4 : ereport(LOG,
6161 : : errmsg("subscription \"%s\" has been disabled because of an error",
6162 : : MySubscription->name));
6163 : :
6164 : : /*
6165 : : * Skip the track_commit_timestamp check when disabling the worker due to
6166 : : * an error, as verifying commit timestamps is unnecessary in this
6167 : : * context.
6168 : : */
6169 : 4 : CheckSubDeadTupleRetention(false, true, WARNING,
6170 : 4 : MySubscription->retaindeadtuples,
6171 : 4 : MySubscription->retentionactive, false);
6172 : :
6173 : 4 : proc_exit(0);
6174 : : }
6175 : :
6176 : : /*
6177 : : * Is current process a logical replication worker?
6178 : : */
6179 : : bool
6180 : 2819 : IsLogicalWorker(void)
6181 : : {
6182 : 2819 : return MyLogicalRepWorker != NULL;
6183 : : }
6184 : :
6185 : : /*
6186 : : * Is current process a logical replication parallel apply worker?
6187 : : */
6188 : : bool
6189 : 2023 : IsLogicalParallelApplyWorker(void)
6190 : : {
6191 [ + + + - ]: 2023 : return IsLogicalWorker() && am_parallel_apply_worker();
6192 : : }
6193 : :
6194 : : /*
6195 : : * Start skipping changes of the transaction if the given LSN matches the
6196 : : * LSN specified by subscription's skiplsn.
6197 : : */
6198 : : static void
6199 : 593 : maybe_start_skipping_changes(XLogRecPtr finish_lsn)
6200 : : {
6201 : : Assert(!is_skipping_changes());
6202 : : Assert(!in_remote_transaction);
6203 : : Assert(!in_streamed_transaction);
6204 : :
6205 : : /*
6206 : : * Quick return if it's not requested to skip this transaction. This
6207 : : * function is called for every remote transaction and we assume that
6208 : : * skipping the transaction is not used often.
6209 : : */
6210 [ + + - + : 593 : if (likely(!XLogRecPtrIsValid(MySubscription->skiplsn) ||
+ + ]
6211 : : MySubscription->skiplsn != finish_lsn))
6212 : 590 : return;
6213 : :
6214 : : /* Start skipping all changes of this transaction */
6215 : 3 : skip_xact_finish_lsn = finish_lsn;
6216 : :
6217 [ + - ]: 3 : ereport(LOG,
6218 : : errmsg("logical replication starts skipping transaction at LSN %X/%08X",
6219 : : LSN_FORMAT_ARGS(skip_xact_finish_lsn)));
6220 : : }
6221 : :
6222 : : /*
6223 : : * Stop skipping changes by resetting skip_xact_finish_lsn if enabled.
6224 : : */
6225 : : static void
6226 : 30 : stop_skipping_changes(void)
6227 : : {
6228 [ + + ]: 30 : if (!is_skipping_changes())
6229 : 27 : return;
6230 : :
6231 [ + - ]: 3 : ereport(LOG,
6232 : : errmsg("logical replication completed skipping transaction at LSN %X/%08X",
6233 : : LSN_FORMAT_ARGS(skip_xact_finish_lsn)));
6234 : :
6235 : : /* Stop skipping changes */
6236 : 3 : skip_xact_finish_lsn = InvalidXLogRecPtr;
6237 : : }
6238 : :
6239 : : /*
6240 : : * Clear subskiplsn of pg_subscription catalog.
6241 : : *
6242 : : * finish_lsn is the transaction's finish LSN that is used to check if the
6243 : : * subskiplsn matches it. If not matched, we raise a warning when clearing the
6244 : : * subskiplsn in order to inform users for cases e.g., where the user mistakenly
6245 : : * specified the wrong subskiplsn.
6246 : : */
6247 : : static void
6248 : 554 : clear_subscription_skip_lsn(XLogRecPtr finish_lsn)
6249 : : {
6250 : : Relation rel;
6251 : : Form_pg_subscription subform;
6252 : : HeapTuple tup;
6253 : 554 : XLogRecPtr myskiplsn = MySubscription->skiplsn;
6254 : 554 : bool started_tx = false;
6255 : :
6256 [ + + - + ]: 554 : if (likely(!XLogRecPtrIsValid(myskiplsn)) || am_parallel_apply_worker())
6257 : 551 : return;
6258 : :
6259 [ + + ]: 3 : if (!IsTransactionState())
6260 : : {
6261 : 1 : StartTransactionCommand();
6262 : 1 : started_tx = true;
6263 : : }
6264 : :
6265 : : /*
6266 : : * Updating pg_subscription might involve TOAST table access, so ensure we
6267 : : * have a valid snapshot.
6268 : : */
6269 : 3 : PushActiveSnapshot(GetTransactionSnapshot());
6270 : :
6271 : : /*
6272 : : * Protect subskiplsn of pg_subscription from being concurrently updated
6273 : : * while clearing it.
6274 : : */
6275 : 3 : LockSharedObject(SubscriptionRelationId, MySubscription->oid, 0,
6276 : : AccessShareLock);
6277 : :
6278 : 3 : rel = table_open(SubscriptionRelationId, RowExclusiveLock);
6279 : :
6280 : : /* Fetch the existing tuple. */
6281 : 3 : tup = SearchSysCacheCopy1(SUBSCRIPTIONOID,
6282 : : ObjectIdGetDatum(MySubscription->oid));
6283 : :
6284 [ - + ]: 3 : if (!HeapTupleIsValid(tup))
6285 [ # # ]: 0 : elog(ERROR, "subscription \"%s\" does not exist", MySubscription->name);
6286 : :
6287 : 3 : subform = (Form_pg_subscription) GETSTRUCT(tup);
6288 : :
6289 : : /*
6290 : : * Clear the subskiplsn. If the user has already changed subskiplsn before
6291 : : * clearing it we don't update the catalog and the replication origin
6292 : : * state won't get advanced. So in the worst case, if the server crashes
6293 : : * before sending an acknowledgment of the flush position the transaction
6294 : : * will be sent again and the user needs to set subskiplsn again. We can
6295 : : * reduce the possibility by logging a replication origin WAL record to
6296 : : * advance the origin LSN instead but there is no way to advance the
6297 : : * origin timestamp and it doesn't seem to be worth doing anything about
6298 : : * it since it's a very rare case.
6299 : : */
6300 [ + - ]: 3 : if (subform->subskiplsn == myskiplsn)
6301 : : {
6302 : : bool nulls[Natts_pg_subscription];
6303 : : bool replaces[Natts_pg_subscription];
6304 : : Datum values[Natts_pg_subscription];
6305 : :
6306 : 3 : memset(values, 0, sizeof(values));
6307 : 3 : memset(nulls, false, sizeof(nulls));
6308 : 3 : memset(replaces, false, sizeof(replaces));
6309 : :
6310 : : /* reset subskiplsn */
6311 : 3 : values[Anum_pg_subscription_subskiplsn - 1] = LSNGetDatum(InvalidXLogRecPtr);
6312 : 3 : replaces[Anum_pg_subscription_subskiplsn - 1] = true;
6313 : :
6314 : 3 : tup = heap_modify_tuple(tup, RelationGetDescr(rel), values, nulls,
6315 : : replaces);
6316 : 3 : CatalogTupleUpdate(rel, &tup->t_self, tup);
6317 : :
6318 [ - + ]: 3 : if (myskiplsn != finish_lsn)
6319 [ # # ]: 0 : ereport(WARNING,
6320 : : errmsg("skip-LSN of subscription \"%s\" cleared", MySubscription->name),
6321 : : errdetail("Remote transaction's finish WAL location (LSN) %X/%08X did not match skip-LSN %X/%08X.",
6322 : : LSN_FORMAT_ARGS(finish_lsn),
6323 : : LSN_FORMAT_ARGS(myskiplsn)));
6324 : : }
6325 : :
6326 : 3 : heap_freetuple(tup);
6327 : 3 : table_close(rel, NoLock);
6328 : :
6329 : 3 : PopActiveSnapshot();
6330 : :
6331 [ + + ]: 3 : if (started_tx)
6332 : 1 : CommitTransactionCommand();
6333 : : }
6334 : :
6335 : : /* Error callback to give more context info about the change being applied */
6336 : : void
6337 : 12147 : apply_error_callback(void *arg)
6338 : : {
6339 : 12147 : ApplyRemoteCtx *ctx = &remote_ctx;
6340 : :
6341 [ + + ]: 12147 : if (ctx->command == 0)
6342 : 11692 : return;
6343 : :
6344 : : Assert(ctx->origin_name);
6345 : :
6346 [ + + ]: 455 : if (ctx->rel == NULL)
6347 : : {
6348 [ - + ]: 338 : if (!TransactionIdIsValid(ctx->remote_xid))
6349 : 0 : errcontext("processing remote data for replication origin \"%s\" during message type \"%s\"",
6350 : : ctx->origin_name,
6351 : : logicalrep_message_type(ctx->command));
6352 [ + + ]: 338 : else if (!XLogRecPtrIsValid(ctx->finish_lsn))
6353 : 264 : errcontext("processing remote data for replication origin \"%s\" during message type \"%s\" in transaction %u",
6354 : : ctx->origin_name,
6355 : : logicalrep_message_type(ctx->command),
6356 : : ctx->remote_xid);
6357 : : else
6358 : 148 : errcontext("processing remote data for replication origin \"%s\" during message type \"%s\" in transaction %u, finished at %X/%08X",
6359 : : ctx->origin_name,
6360 : : logicalrep_message_type(ctx->command),
6361 : : ctx->remote_xid,
6362 : 74 : LSN_FORMAT_ARGS(ctx->finish_lsn));
6363 : : }
6364 : : else
6365 : : {
6366 [ + - ]: 117 : if (ctx->remote_attnum < 0)
6367 : : {
6368 [ + + ]: 117 : if (!XLogRecPtrIsValid(ctx->finish_lsn))
6369 : 4 : errcontext("processing remote data for replication origin \"%s\" during message type \"%s\" for replication target relation \"%s.%s\" in transaction %u",
6370 : : ctx->origin_name,
6371 : : logicalrep_message_type(ctx->command),
6372 : 2 : ctx->rel->remoterel.nspname,
6373 : 2 : ctx->rel->remoterel.relname,
6374 : : ctx->remote_xid);
6375 : : else
6376 : 230 : errcontext("processing remote data for replication origin \"%s\" during message type \"%s\" for replication target relation \"%s.%s\" in transaction %u, finished at %X/%08X",
6377 : : ctx->origin_name,
6378 : : logicalrep_message_type(ctx->command),
6379 : 115 : ctx->rel->remoterel.nspname,
6380 : 115 : ctx->rel->remoterel.relname,
6381 : : ctx->remote_xid,
6382 : 115 : LSN_FORMAT_ARGS(ctx->finish_lsn));
6383 : : }
6384 : : else
6385 : : {
6386 [ # # ]: 0 : if (!XLogRecPtrIsValid(ctx->finish_lsn))
6387 : 0 : errcontext("processing remote data for replication origin \"%s\" during message type \"%s\" for replication target relation \"%s.%s\" column \"%s\" in transaction %u",
6388 : : ctx->origin_name,
6389 : : logicalrep_message_type(ctx->command),
6390 : 0 : ctx->rel->remoterel.nspname,
6391 : 0 : ctx->rel->remoterel.relname,
6392 : 0 : ctx->rel->remoterel.attnames[ctx->remote_attnum],
6393 : : ctx->remote_xid);
6394 : : else
6395 : 0 : errcontext("processing remote data for replication origin \"%s\" during message type \"%s\" for replication target relation \"%s.%s\" column \"%s\" in transaction %u, finished at %X/%08X",
6396 : : ctx->origin_name,
6397 : : logicalrep_message_type(ctx->command),
6398 : 0 : ctx->rel->remoterel.nspname,
6399 : 0 : ctx->rel->remoterel.relname,
6400 : 0 : ctx->rel->remoterel.attnames[ctx->remote_attnum],
6401 : : ctx->remote_xid,
6402 : 0 : LSN_FORMAT_ARGS(ctx->finish_lsn));
6403 : : }
6404 : : }
6405 : : }
6406 : :
6407 : : /*
6408 : : * Set information identifying the remote transaction currently being
6409 : : * applied, kept for the duration of that transaction.
6410 : : *
6411 : : * This must be called for every message type that begins or resumes applying
6412 : : * a remote transaction's changes (BEGIN, BEGIN PREPARE, STREAM START, STREAM
6413 : : * COMMIT, STREAM PREPARE), since interleaved transactions (possible only for
6414 : : * streaming) would otherwise leave stale values from whichever transaction
6415 : : * last called this.
6416 : : *
6417 : : * Callers normally pass the top-level transaction's xid. The exception is a
6418 : : * STREAM ABORT, which passes the xid of the (sub)transaction being aborted so
6419 : : * that the error context names whatever failed; that is a subxid only when a
6420 : : * subtransaction rolls back, and the top-level xid otherwise. Nothing else
6421 : : * observes a subxid recorded this way, because no change is applied between a
6422 : : * STREAM ABORT and the STREAM START or STREAM COMMIT/PREPARE that follows it,
6423 : : * and each of those calls this again with the transaction's own values.
6424 : : */
6425 : : static inline void
6426 : 2999 : set_remote_transaction_info(TransactionId xid, XLogRecPtr lsn)
6427 : : {
6428 : 2999 : remote_ctx.remote_xid = xid;
6429 : 2999 : remote_ctx.finish_lsn = lsn;
6430 : 2999 : }
6431 : :
6432 : : /* Reset all information of the remote transaction context */
6433 : : static inline void
6434 : 1450 : reset_apply_remote_context(void)
6435 : : {
6436 : 1450 : remote_ctx.command = 0;
6437 : 1450 : remote_ctx.rel = NULL;
6438 : 1450 : remote_ctx.remote_attnum = -1;
6439 : 1450 : set_remote_transaction_info(InvalidTransactionId, InvalidXLogRecPtr);
6440 : 1450 : }
6441 : :
6442 : : /*
6443 : : * Request wakeup of the workers for the given subscription OID
6444 : : * at commit of the current transaction.
6445 : : *
6446 : : * This is used to ensure that the workers process assorted changes
6447 : : * as soon as possible.
6448 : : */
6449 : : void
6450 : 399 : LogicalRepWorkersWakeupAtCommit(Oid subid)
6451 : : {
6452 : : MemoryContext oldcxt;
6453 : :
6454 : 399 : oldcxt = MemoryContextSwitchTo(TopTransactionContext);
6455 : 399 : on_commit_wakeup_workers_subids =
6456 : 399 : list_append_unique_oid(on_commit_wakeup_workers_subids, subid);
6457 : 399 : MemoryContextSwitchTo(oldcxt);
6458 : 399 : }
6459 : :
6460 : : /*
6461 : : * Wake up the workers of any subscriptions that were changed in this xact.
6462 : : */
6463 : : void
6464 : 634417 : AtEOXact_LogicalRepWorkers(bool isCommit)
6465 : : {
6466 [ + + + + ]: 634417 : if (isCommit && on_commit_wakeup_workers_subids != NIL)
6467 : : {
6468 : : ListCell *lc;
6469 : :
6470 : 385 : LWLockAcquire(LogicalRepWorkerLock, LW_SHARED);
6471 [ + - + + : 770 : foreach(lc, on_commit_wakeup_workers_subids)
+ + ]
6472 : : {
6473 : 385 : Oid subid = lfirst_oid(lc);
6474 : : List *workers;
6475 : : ListCell *lc2;
6476 : :
6477 : 385 : workers = logicalrep_workers_find(subid, true, false);
6478 [ + + + + : 470 : foreach(lc2, workers)
+ + ]
6479 : : {
6480 : 85 : LogicalRepWorker *worker = (LogicalRepWorker *) lfirst(lc2);
6481 : :
6482 : 85 : logicalrep_worker_wakeup_ptr(worker);
6483 : : }
6484 : : }
6485 : 385 : LWLockRelease(LogicalRepWorkerLock);
6486 : : }
6487 : :
6488 : : /* The List storage will be reclaimed automatically in xact cleanup. */
6489 : 634417 : on_commit_wakeup_workers_subids = NIL;
6490 : 634417 : }
6491 : :
6492 : : /*
6493 : : * Allocate the origin name in long-lived context for error context message.
6494 : : */
6495 : : void
6496 : 508 : set_apply_error_context_origin(char *originname)
6497 : : {
6498 : 508 : remote_ctx.origin_name = MemoryContextStrdup(ApplyContext, originname);
6499 : 508 : }
6500 : :
6501 : : /*
6502 : : * Return the action to be taken for the given transaction. See
6503 : : * TransApplyAction for information on each of the actions.
6504 : : *
6505 : : * *winfo is assigned to the destination parallel worker info when the leader
6506 : : * apply worker has to pass all the transaction's changes to the parallel
6507 : : * apply worker.
6508 : : */
6509 : : static TransApplyAction
6510 : 347251 : get_transaction_apply_action(TransactionId xid, ParallelApplyWorkerInfo **winfo)
6511 : : {
6512 : 347251 : *winfo = NULL;
6513 : :
6514 [ + + ]: 347251 : if (am_parallel_apply_worker())
6515 : : {
6516 : 69135 : return TRANS_PARALLEL_APPLY;
6517 : : }
6518 : :
6519 : : /*
6520 : : * If we are processing this transaction using a parallel apply worker
6521 : : * then either we send the changes to the parallel worker or if the worker
6522 : : * is busy then serialize the changes to the file which will later be
6523 : : * processed by the parallel worker.
6524 : : */
6525 : 278116 : *winfo = pa_find_worker(xid);
6526 : :
6527 [ + + + + ]: 278116 : if (*winfo && (*winfo)->serialize_changes)
6528 : : {
6529 : 5037 : return TRANS_LEADER_PARTIAL_SERIALIZE;
6530 : : }
6531 [ + + ]: 273079 : else if (*winfo)
6532 : : {
6533 : 68906 : return TRANS_LEADER_SEND_TO_PARALLEL;
6534 : : }
6535 : :
6536 : : /*
6537 : : * If there is no parallel worker involved to process this transaction
6538 : : * then we either directly apply the change or serialize it to a file
6539 : : * which will later be applied when the transaction finish message is
6540 : : * processed.
6541 : : */
6542 [ + + ]: 204173 : else if (in_streamed_transaction)
6543 : : {
6544 : 103197 : return TRANS_LEADER_SERIALIZE;
6545 : : }
6546 : : else
6547 : : {
6548 : 100976 : return TRANS_LEADER_APPLY;
6549 : : }
6550 : : }
|