Skip to content

Host-capture Re_size in viscous and non-Newtonian device code - #2001

Draft
sbryngelson wants to merge 1 commit into
MFlowCode:masterfrom
sbryngelson:fix/amdflang-re-size-stale
Draft

sbryngelson wants to merge 1 commit into
MFlowCode:masterfrom
sbryngelson:fix/amdflang-re-size-stale

Conversation

@sbryngelson

@sbryngelson sbryngelson commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

#1588 worked around amdflang reading the declare-target Re_size stale
(as 0) across translation units by copying it into firstprivate locals
in the Riemann solvers. Device code in other modules still read the
module variable directly, so on AMD flang it could silently drop
viscosity the same way:

  • s_compute_axis_inv_re (m_viscous), called from the four cylindrical
    axis stress kernels
  • s_compute_mixture_inv_re (m_hb_function), called from the axis
    routine and from the cylindrical and Cartesian viscous source-flux
    kernels in m_riemann_state

Apply the same pattern: each enclosing kernel takes Re_size_loc1/2 as
firstprivate host copies and passes them down, and the device routines
use merge(Re_size_loc1, Re_size_loc2, i == 1) in place of Re_size(i),
as s_compute_interface_reynolds already does.

Values are identical to Re_size, so results do not change on CPU or
NVIDIA: viscous, non-Newtonian and cylindrical tests pass unchanged and
the NVHPC OpenACC build compiles. Not tested on AMD hardware.

MFlowCode#1588 worked around amdflang reading the declare-target Re_size stale
(as 0) across translation units by copying it into firstprivate locals
in the Riemann solvers. Device code in other modules still read the
module variable directly, so on AMD flang it could silently drop
viscosity the same way:

- s_compute_axis_inv_re (m_viscous), called from the four cylindrical
  axis stress kernels
- s_compute_mixture_inv_re (m_hb_function), called from the axis
  routine and from the cylindrical and Cartesian viscous source-flux
  kernels in m_riemann_state

Apply the same pattern: each enclosing kernel takes Re_size_loc1/2 as
firstprivate host copies and passes them down, and the device routines
use merge(Re_size_loc1, Re_size_loc2, i == 1) in place of Re_size(i),
as s_compute_interface_reynolds already does.

Values are identical to Re_size, so results do not change on CPU or
NVIDIA: viscous, non-Newtonian and cylindrical tests pass unchanged and
the NVHPC OpenACC build compiles. Not tested on AMD hardware.
@github-actions

Copy link
Copy Markdown

Lines of Code

File Lines Diff
src/simulation/m_viscous.fpp 1028 +7
src/simulation/m_riemann_state.fpp 1189 +4
src/simulation/m_hb_function.fpp 62 +1
Directory Lines Diff
simulation 28457 +12
total 47438 +12

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant