MethodologyCitableFor research partners

Research Methodology: The Shared Design, Data Boundaries and Collaboration Terms Behind Eight Angling-App Studies

Angling-app catch records are a strange kind of data: large, finely resolved in space and time, and recording only success. Over the past decade several large fishing apps in Europe and North America have brought such data into the fisheries literature, mainly as a supplement for species distribution and relative abundance. We took a different route: treating the data as a testing ground for behaviour and lore. Which of the things anglers believe survive once effort is put back in the denominator?

Eight studies in, a fixed practice has formed. This page writes it down so that a researcher citing any one of them knows the data's boundaries, how confounding is controlled, and what this data can never do.

Version 2026-09-05 · Data 2025-01-01 to 2026-09-04 · 中文版

8
hypothesis-testing studies (Sept 2026)
3
null results with an excluded magnitude
10
shared method principles
100%
with a negative control and reported power
Data source and scope

Catches voluntarily logged by DudeFishinG app users. Each record carries species (AI-identified or angler-chosen, reviewed), time (photo EXIF or manual), location (GPS, EXIF or manual, mapped to a water-body node classified freshwater / marine / estuary), optional length, released or kept, method, and weather, pressure, wave height, tide and sun altitude attached from the nearest station or forecast at logging time. Period 2025-01-01 to 2026-09-04; non-fish, test accounts and anomalies listed in the cleaning report are excluded (see the verified-catch cleaning report and the data methodology page).

What is absent matters as much: no blank trips, no trip start or end, no bait or hours. So there is no catch per unit effort, and every study is designed around that gap. This page and every study publish only de-identified aggregates, figures and methods; no individual record, account or coordinate is released.

The eight studies
QuestionDesignControlResult
Do moon phase and tidal range change catch rates?No effect (magnitude excluded)Daily catch and angler series with weekday and month fixed effects, harmonic regression on lunar age and tidal range; circular-shift permutation inference.Freshwater versus marine; night and day split.No detectable effect; amplitudes of 1.15× or more (1.28× at night) are excluded.
Do anglers keep invasive fish rather than release them?Question reframedAngler-fixed-effect logistic regression of release on species status, method and environment; within-angler contrasts.Natives versus invasives; stratified by method.Release (87%) is method culture, almost unrelated to species status; only giant snakehead and sailfin catfish are exceptions.
Is the angler movement network a pathway for invasive spread?No effect (magnitude excluded)Water-body network from shared anglers; connectivity predicting invasive presence; edge-shuffle permutation.The same metric applied to native species.Connectivity predicts invasives and natives equally (0.50 vs 0.58): the network maps angler traffic, not spread.
Does the fishing method select the species, or does the place?Question reframedMutual information between species and method, conditional on water body; within-angler contrasts; 5,021 water-body fixed effects absorbed by alternating demeaning.Comparisons within angler and within water body.75% of apparent method selectivity is place; gear reliably changes size (pole 20% smaller than lure), not species.
Do fish bite hard when pressure drops before a front?Partly supported, smallA-priori power analysis; daily catch per angler on 24-hour pressure change with FWL-residualised fixed effects; circular-shift permutation; jackknife.Freshwater versus marine; EXIF subset.A real but small signal: freshwater +13% on drop days, marine −8%, 7% more anglers out; no 'feeding frenzy' magnitude.
Are dawn and dusk really better than midday?No effect (magnitude excluded)Within-trip exposure correction: sun altitude every 5 minutes between first and last catch; uniform within-window redistribution as the constant-rate null.Null from within-window permutation; EXIF subset.While fishing, twilight is no better than midday; the raw peak is when people fish; a 1.5× edge would be detected 100% of the time.
Do typhoon warnings change angler behaviour, and which warning works?SupportedNatural experiment: official sea/land warning windows against daily logging anglers; same-weekday ±4-week expectation; independent per-typhoon window shifts as the null.Freshwater anglers (same holidays and rain, no sea state).Land warnings cut marine participation to 0.40×; sea-only 0.85× (indistinguishable from none); no rebound after the lift.
How many caught fish are juveniles, and what do anglers release by?SupportedLengths against FishBase / peer-reviewed L50; angler bootstrap; species-fixed-effect logit; within-angler within-species pairs; discontinuity test at L50.Freshwater invasives (release should not depend on size).66% of marine natives are immature; release is size-selective (odds 3.7× per halving) for natives only; no jump at L50.

Four kinds of outcome: no effect (with the excluded magnitude), supported, partly supported (real but small), question reframed (an effect exists but belongs to another variable).

Shared method principles
  1. Hypotheses fixed before looking. Each study lists H1–H4 with direction and threshold, Bonferroni for multiplicity; outcomes are supported / rejected / not supported / not established, never re-framed after the fact.
  2. Put effort back in the denominator. Anglers do not log blank trips, so every 'more catches at X' is first asked whether X is simply when more people fish. Denominators are angler counts, places, or within-trip exposure.
  3. A negative control in every study. Freshwater versus marine, natives versus invasives, the same metric on a group where no effect should exist. An effect that appears in the control is measuring something else.
  4. Inference by permutation and natural experiment. Circular shifts, independent event-window shifts, within-window redistribution: the null comes from the data's own structure, not parametric assumptions; exogenous events (typhoon warnings) provide quasi-experiments.
  5. Within-unit designs. Same angler, same water body, same trip, same species: paired contrasts absorb skill, place and habit at once.
  6. Power and minimum detectable effect always reported. A null is only meaningful with 'a true X× effect would have been detected'; the excluded effect size is part of the conclusion.
  7. Robustness subsets. EXIF-timestamp subset (reliable clock), ruler-measured subset, catch counts versus angler counts, leave-one-event-out jackknife; the main conclusion must hold direction across subsets.
  8. Measured and inferred kept apart. One section per page: measured numbers and intervals above, interpretation, assumptions and bias direction below. Readers may trust only the top half.
  9. One source of truth for every number. Every figure on a page is generated from the analysis results JSON into TypeScript constants shared by both languages, the citation files and the dataset catalog; nothing is hand-edited.
  10. Read-only analysis, aggregates only. Analyses never modify production data; what is published is aggregates, figures and methods; raw records and coordinates are not released.
Known data-quality issues

These are not hidden weaknesses; they are the starting point of every design. Read this section before citing any figure.

  • Only catches are logged. Blank trips are invisible, so there is no catch per unit effort, only 'anglers logging a catch' as a participation proxy.
  • Night heaping of manually entered times: fished by day, logged at night with the default 'now'. Every intraday analysis carries an EXIF-timestamp subset. (first measured in Are dawn and dusk really better than midday?)
  • 58% of lengths are multiples of 5 cm, 47% even among ruler-flagged records; most lengths are estimated. Jitter sensitivity confirms conclusions are unaffected. (first measured in How many caught fish are juveniles, and what do anglers release by?)
  • Freshwater / marine is assigned from the water-body classification; estuaries are brackish and grouped with marine. Misclassification dilutes the contrast between groups.
  • The user base grew several-fold from 2025 to 2026; every temporal comparison uses local expectations (same weekday ±4 weeks) or fixed effects, never a whole-period mean. (first measured in Do typhoon warnings change angler behaviour, and which warning works?)
  • Length, release flag and photo are all optional; records carrying them skew toward memorable fish. Subset and full-sample results are shown side by side.
  • Weather and wave height attached to a catch come from the nearest station or forecast at logging time, not on-site measurement. (first measured in Do fish bite hard when pressure drops before a front?)
What this data can and cannot do

Can

  • Relative comparisons with a control (marine vs freshwater, native vs invasive, warning day vs normal)
  • Participation: daily anglers logging a catch, and its response to exogenous events
  • Within-trip relative bite rate (no effort reporting needed)
  • Structure and selectivity across species, place, method and size
  • Angler behaviour: release decisions, size thresholds, going out under risk

Cannot (pending more data or external data)

  • Catch per unit effort or absolute catch rates
  • Abundance, stock status or biomass trends
  • Blank trips, trip hours, bait use
  • Causal claims without a natural experiment or control

The "cannot" column is not permanent. Trip units (start and end fishing) are being piloted in the app; once blank trips and hours accumulate over enough seasons, CPUE and abundance trends become possible, and this page will be reissued.

How a study is produced
  1. Topic. Prefer questions with lore behind them, effort confounding, and a feasible control; run power first and drop topics that cannot be detected.
  2. Fix hypotheses. H1–H4 with direction, threshold, control and multiplicity correction, written in the analysis script header.
  3. Read-only extract. Export only the needed fields from production; write nothing back.
  4. Analyse. Results JSON saved section by section; 1,000–3,000 permutations; robustness subsets run alongside.
  5. Generate pages. Results JSON → TypeScript constants (single source of truth) → both language pages, charts, BibTeX / CSL-JSON, dataset catalog; every number from one source.
  6. Measured / inferred. One section per page; in-app product ideas are listed as suggestions only and never change the app as part of a study.
  7. Archive. Analysis date, data period, code and results are versioned; pages do not update silently.
Citation and collaborationCollaboration welcome

Every study page has BibTeX and CSL-JSON; please retain data period, sample size and analysis date. This page is citable too: BibTeX · CSL-JSON · dataset catalog. Public pages are free to read and cite page by page. Raw records are not available for download, and bulk reuse licensing for aggregated data is still under review.

What we can provide: analysis code and bin definitions for each study; the aggregate tables named in each Methods section (day-level observed/expected, per-species length bins, sun-altitude exposure bins); joint recomputation and co-authorship with your external data.

What we most want: length at maturity for Taiwan populations; coast guard rescue or incident records (to link with the typhoon-warning study); timestamped catch records from other regions or communities for cross-regional replication of the within-trip and negative-control designs. Write to support@dudefishing.net.

Frequently asked questions

What is the data?

Catches voluntarily logged by DudeFishinG app users in Taiwan: species, time, location (mapped to a water body), length, released or kept, method, plus weather and tide attached at logging time. Period 2025-01-01 to 2026-09-04, about thirty thousand valid fish catches by several thousand anglers. About four in ten records carry a photo EXIF timestamp. There are no blank trips and no trip durations.

How can conclusions be drawn without effort data?

Three ways. Replace the denominator with something measurable: daily anglers logging a catch (participation), places, or the exposure time inside a trip between first and last catch. Make only relative comparisons so a control group carries the same bias. Use exogenous events such as typhoon warnings as natural experiments. The data cannot yield catch per unit effort, but it can answer whether X changed behaviour or a relative rate.

Why are half the studies null results?

Because most angling lore is an artefact of effort confounding: dawn and dusk, moon phase, fronts, method selectivity all flatten or shrink once effort is in the denominator. A null here is not 'nothing found' but 'an effect of X× or more would have been detected and was not', and every study reports its minimum detectable effect.

Can researchers obtain the raw data?

No; raw records and coordinates are not released. Available are all published aggregates and figures (free to cite, with BibTeX and CSL-JSON), analysis code and bin definitions on request, and the aggregate tables named in each Methods section (day-level observed/expected, per-species length bins). If you hold linkable external data (incident records, local maturity lengths, timestamped catch records from other regions) we will recompute jointly and co-author. Write to support@dudefishing.net.

Will the numbers change?

Each study has a fixed data period and analysis date, and every figure is generated from the results file, so pages do not update silently. If a study is redone with more data it is released under a new analysis date and the previous version stays linked.

What comes next?

Questions that need trip units, blank-trip records or longer series (true CPUE, abundance trends, climate inter-annual effects) are on hold until data accumulates. What can be done now is cross-regional replication: the within-trip and negative-control designs apply to any timestamped catch dataset.