Requesting download of a log which is not present on the card - a
hole in the log sequence, advertised in the log list as a
zero-size/zero-time entry - set _open_error_ms when the read-side
open failed. That flag pauses logging, fails downloads of logs
which do exist with an immediate zero-length EOF and fails arming
checks (PreArm: Logging failed) for the next five seconds.
A missing log is an expected condition, unlike the I/O errors from a
failing card the five-second backoff protects against, so do not set
the flag for ENOENT.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
_open_error_ms is set whenever an open fails - on the read path as well
as the write path - but what it exists to gate is all on the writer's
side: starting a log, writing to one, and reporting whether logging
works. start_new_log() sets it speculatively too, before doing
anything else, so that a GCS_SEND_TEXT() further down that path cannot
recurse back into opening a log; it is cleared only once the new file
is actually open. Every logging restart therefore leaves
recent_open_error() true for up to LOGGER_FILE_REOPEN_MS - five
seconds - with nothing having gone wrong.
get_log_data() consulted that flag, and start_new_log() closes _read_fd
too, so a download straddling a logging restart reopens the file inside
the window the restart itself created and gets -1 back. The MAVLink
log transfer reports -1 as EOF, so the client is handed a zero-length
LOG_DATA in the middle of the file and stops there:
REQ id=7 ofs=56259 count=67 -> DAT ofs=56259 count=67
REQ id=7 ofs=56211 count=48 -> DAT ofs=56211 count=0
Whether the writer can open a file has no bearing on reading one. Drop
the consultation; if the filesystem is genuinely struggling then the
open in the read path fails and returns -1 by itself.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A writer thread which slipped a block into the write buffer between
_write_fd becoming valid and the buffer being cleared would have those
bytes silently discarded. If those bytes were a FMT message the
backend also marks that format as written-to-this-log, so the format
is never re-emitted: the entire log then contains records of that type
with no FMT to decode them by. Logs exhibiting exactly this (PARM,
BARO and MAG records with no corresponding FMT anywhere in the file)
have been captured from parallel autotest runs.
Clear the buffer before the fd becomes visible so nothing can enter
the buffer and be discarded.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
We used to produce files hwih looked like log1.BIN. We moved to 00000001.BIN instead so things collate.
This code allowed the autopilot to return data from SD cards which had old logs on them.
this sets the logging rate max when disarmed. In combination with
LOG_DISARMED=3 it gives a very nice setup to get always on logging
with very little addition to the log sizes. It is particularly useful
in combination with LOG_REPLAY=1
when LOG_DISARMED is set to 3 then we log while disarmed but if we
reboot without ever arming the log is discarded. This allows for using
LOG_DISARMED without filling the microSD.
this fixes two issues:
The first issue that if we are missing a log file in the middle of the
list then it was not possible to download recent logs, as we get the
incorrect value for total number of logs. This happened for me with
107 logs, with log62 missing from the microSD. It would only show 45
available logs, so the most recent logs could not be downloaded.
The second issue is that get_num_logs() was very slow if there were a
lot of log files in a directory. This would cause EKF errors and ESC
resets. Using a opendir/readdir loop is much faster (approx 10x faster
in my testing with 107 logs on a MatekH743).
this fixes a problem with sdcards where file open is very slow. It can
trigger a watchdog if it is slow enough. Peter and I hit this issue on
a pixracer today with a new sd card
this fixes an issue where the sd card fails in flight and then
re-mounts. When that happens the logging backend can trigger a new log
open. That causes filesystem operations in the main thread while
flying. That can cause long delays or even a watchdog.
Thanks to Giacomo for noticing this on his flying wing