Martin Kroeker [Mon, 13 Apr 2020 10:18:48 +0000 (12:18 +0200)]
Add (empty) read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:17:41 +0000 (12:17 +0200)]
Add (empty) read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:16:44 +0000 (12:16 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:14:58 +0000 (12:14 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:14:06 +0000 (12:14 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:12:54 +0000 (12:12 +0200)]
Add (empty) read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:11:58 +0000 (12:11 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:10:37 +0000 (12:10 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:09:24 +0000 (12:09 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:06:40 +0000 (12:06 +0200)]
Merge pull request #46 from xianyi/develop
rebase
Martin Kroeker [Sun, 12 Apr 2020 20:34:41 +0000 (22:34 +0200)]
Merge pull request #2551 from martin-frbg/issue2538-2
Increase BUFFER_SIZEs and add a safeguard; supply GEMM_R for POWER8/9
Martin Kroeker [Sun, 12 Apr 2020 17:47:02 +0000 (19:47 +0200)]
Fix parameter overflow
Martin Kroeker [Sun, 12 Apr 2020 17:45:36 +0000 (19:45 +0200)]
Add safeguards for sufficient BUFFER_SIZE
Martin Kroeker [Sun, 12 Apr 2020 17:44:48 +0000 (19:44 +0200)]
Increase default BUFFER_SIZE on ARM, ZARCH and newer x86_64, add GEMM_R for POWER8/9
As shown in #2538, default buffersizes on some platforms were smaller than required in memory.c
and the requirement could never be fulfilled for a calculated GEMM_R on PPC given the fomula used
Martin Kroeker [Sun, 12 Apr 2020 17:39:05 +0000 (19:39 +0200)]
Merge pull request #45 from xianyi/develop
rebase
Martin Kroeker [Fri, 10 Apr 2020 22:35:38 +0000 (00:35 +0200)]
Merge pull request #2547 from sharvil/develop
Add API to set thread affinity on Linux.
Martin Kroeker [Fri, 10 Apr 2020 22:35:07 +0000 (00:35 +0200)]
Merge pull request #2541 from bapt/develop
libname: treat FreeBSD and DragonFly like linux and sunos
Martin Kroeker [Fri, 10 Apr 2020 22:34:04 +0000 (00:34 +0200)]
Merge pull request #2548 from gxw-loongson/develop
Add a GENERIC target for 64bit MIPS
Martin Kroeker [Fri, 10 Apr 2020 21:54:40 +0000 (23:54 +0200)]
Merge pull request #2549 from martin-frbg/fixthreadtest
Match thread count in cpp_thread_test to host capability
Martin Kroeker [Fri, 10 Apr 2020 20:06:44 +0000 (22:06 +0200)]
Match thread count to machine capability
gxw [Thu, 9 Apr 2020 11:25:13 +0000 (19:25 +0800)]
Fix compilation problem on loongson platform
Using "make TARGET=GENERIC" on loongson platform will get the following
error messages:
"make[1]: *** No rule to make target 'sgemm_incopy.o', needed by 'libs'"
Add kernel/mips64/KERNEL.generic to slove the problem.
Sharvil Nanavati [Wed, 8 Apr 2020 19:47:41 +0000 (12:47 -0700)]
Add API to set thread affinity on Linux.
Issue: #2545
Martin Kroeker [Wed, 8 Apr 2020 09:04:51 +0000 (11:04 +0200)]
Add another memory barrier for ARM and a multicore test run on ThunderX to help detect such issues (#2544)
* Add another memory barrier in memory.c to prevent races in memory slot allocation
* Add an all-core test on Drone.io's ThunderX platform and modify dgemm_tester to use all 96 cores
Martin Kroeker [Sat, 4 Apr 2020 20:48:53 +0000 (22:48 +0200)]
Merge pull request #44 from xianyi/develop
Add a Z13 build to the Travis configuration (#2542)
Martin Kroeker [Sat, 4 Apr 2020 20:46:58 +0000 (22:46 +0200)]
Merge pull request #43 from martin-frbg/revert-42-z12ci
Revert 42 z12ci to keep forked develop clean
Martin Kroeker [Sat, 4 Apr 2020 20:45:01 +0000 (22:45 +0200)]
Revert "Add IBM Z to Travis configuration (#42)"
This reverts commit
7972beb3754409db9af3c22dbbe7bd8075c09f6e.
Martin Kroeker [Fri, 3 Apr 2020 14:02:11 +0000 (16:02 +0200)]
Add a Z13 build to the Travis configuration (#2542)
* Add IBM Z to Travis configuration
Martin Kroeker [Fri, 3 Apr 2020 13:59:18 +0000 (15:59 +0200)]
Add IBM Z to Travis configuration (#42)
* Add IBM Z to Travis configuration
Baptiste Daroussin [Fri, 3 Apr 2020 04:20:42 +0000 (06:20 +0200)]
libname: treat FreeBSD and DragonFly like linux and sunos
There is no difference in the way libnames are handle between FreeBSD
and linux or sunos. FreeBSD and DragonFly prefers having sonames as well
Martin Kroeker [Thu, 2 Apr 2020 08:32:19 +0000 (10:32 +0200)]
Merge pull request #41 from xianyi/develop
rebase
Martin Kroeker [Thu, 2 Apr 2020 08:30:37 +0000 (10:30 +0200)]
Make ARMV7 compile with xcode and add a CI job for it (#2537)
* Add an ARMV7 iOS build on Travis
* thread_local appears to be unavailable on ARMV7 iOS
* Add no-thumb option for ARMV7 IOS build to get it to accept DMB ISH
* Make local labels in macros of nrm2_vfpv3.S compatible with the xcode assembler
Martin Kroeker [Wed, 1 Apr 2020 18:00:13 +0000 (20:00 +0200)]
Merge pull request #2536 from martin-frbg/recurs
Add "recursive" option for LAPACK builds with ifort or pgfort as well
Martin Kroeker [Wed, 1 Apr 2020 13:39:16 +0000 (15:39 +0200)]
ifort and pgfort need "recursive" for safe compilation of LAPACK as well
Martin Kroeker [Wed, 1 Apr 2020 13:38:07 +0000 (15:38 +0200)]
ifort and pgfort need "recursive" for compiling LAPACK as well
as shown in Reference-LAPACK issue 401 (their PR 403)
Martin Kroeker [Tue, 31 Mar 2020 18:53:13 +0000 (20:53 +0200)]
Merge pull request #2534 from martin-frbg/issue2496
Fix zero initialization for beta=0 case
Martin Kroeker [Tue, 31 Mar 2020 14:53:56 +0000 (16:53 +0200)]
fix initialization to zero in the NEON SGEMM_BETA kernel as well
Martin Kroeker [Mon, 30 Mar 2020 22:21:02 +0000 (00:21 +0200)]
Fix zero initialization for beta=0 case
use immediate initialization instead of multiplication in case register content is a NaN
Martin Kroeker [Mon, 30 Mar 2020 18:15:59 +0000 (20:15 +0200)]
Merge pull request #2520 from wjc404/develop
Fix avx512 sgemm performance bug when ldc is a multiple of 1024
Martin Kroeker [Mon, 30 Mar 2020 18:15:37 +0000 (20:15 +0200)]
Merge pull request #2533 from martin-frbg/gemmdirect2
Use runtime check for AVX512 capability in DYNAMIC_ARCH builds made on SKX
Martin Kroeker [Thu, 26 Mar 2020 20:25:39 +0000 (21:25 +0100)]
Expose the support_avx512 function provided in dynamic.c
Martin Kroeker [Thu, 26 Mar 2020 20:12:56 +0000 (21:12 +0100)]
Use runtime check for AVX512 (sgemm_direct) capability when using DYNAMIC_ARCH
Martin Kroeker [Thu, 26 Mar 2020 20:06:51 +0000 (21:06 +0100)]
Merge pull request #39 from xianyi/develop
rebase
Martin Kroeker [Tue, 24 Mar 2020 14:44:46 +0000 (15:44 +0100)]
Merge pull request #2530 from martin-frbg/dynmsg
Add message highlighting minimum target choice at end of DYNAMIC_ARCH…
Martin Kroeker [Tue, 24 Mar 2020 14:44:27 +0000 (15:44 +0100)]
Merge pull request #2529 from shengyang-3390/dev1
add ctest for drotm and modified ctest for drot.
Martin Kroeker [Mon, 23 Mar 2020 18:35:51 +0000 (19:35 +0100)]
Add message highlighting minimum target choice at end of DYNAMIC_ARCH builds
related to #2526
Martin Kroeker [Mon, 23 Mar 2020 11:47:19 +0000 (12:47 +0100)]
Merge pull request #2527 from martin-frbg/gemmdirect
Avoid calling DIRECT codepath in DYNAMIC_ARCH on non-SKX
shengyang [Sat, 21 Mar 2020 07:58:21 +0000 (15:58 +0800)]
add ctest for drotm and modified ctest for drot.
make sure that test cases cover all code path when kernel uses looping unrolling.
Martin Kroeker [Sun, 22 Mar 2020 13:33:16 +0000 (14:33 +0100)]
Avoid calling DIRECT codepath in DYNAMIC_ARCH on non-SKX
Martin Kroeker [Sun, 22 Mar 2020 00:03:42 +0000 (01:03 +0100)]
Merge pull request #2521 from martin-frbg/cm-avx512
Use proper extension on the avx512 testcase filename
Martin Kroeker [Sat, 21 Mar 2020 17:47:48 +0000 (18:47 +0100)]
Merge pull request #2525 from andreas-schwab/develop
Fix ARCHCONFIG for Neoverse-N1
Andreas Schwab [Sat, 21 Mar 2020 16:33:33 +0000 (17:33 +0100)]
Fix ARCHCONFIG for Neoverse-N1
../config_kernel.h:24:9: warning: missing whitespace after the macro name
24 | #define ARMV8-march armv8.2-a
| ^~~~~
Martin Kroeker [Fri, 20 Mar 2020 22:05:53 +0000 (23:05 +0100)]
Use proper extension on the avx512 testcase filename
The need to call it .tmp existed only when it was generated by a tmpfile call, and the "-x c" option to tell the compiler it is actually a C source is not universally supported (this broke the test with clang-cl at least)
Martin Kroeker [Fri, 20 Mar 2020 22:00:06 +0000 (23:00 +0100)]
Merge pull request #2518 from shengyang-3390/dev
add ctest for srotm and modified ctest for srot.
Martin Kroeker [Fri, 20 Mar 2020 21:57:44 +0000 (22:57 +0100)]
Merge pull request #2519 from martin-frbg/issue2472
Fix cmake compilation with ifort on Windows
wjc404 [Fri, 20 Mar 2020 21:46:18 +0000 (21:46 +0000)]
Update param.h
wjc404 [Fri, 20 Mar 2020 21:42:10 +0000 (05:42 +0800)]
Add files via upload
Martin Kroeker [Fri, 20 Mar 2020 00:08:10 +0000 (01:08 +0100)]
Make ifort on Windows create lowercase symbols with appended underscore
tentative fix for #2472
Martin Kroeker [Fri, 20 Mar 2020 00:05:22 +0000 (01:05 +0100)]
Merge pull request #38 from xianyi/develop
rebase
shengyang [Wed, 18 Mar 2020 06:17:32 +0000 (14:17 +0800)]
add ctest for srotm and modified ctest for srot.
make sure that test cases cover all code path when kernel uses looping unrolling.
Martin Kroeker [Tue, 17 Mar 2020 09:12:53 +0000 (10:12 +0100)]
Merge pull request #2517 from wjc404/develop
Temporary fix for SKX STRSM
wjc404 [Tue, 17 Mar 2020 04:52:55 +0000 (12:52 +0800)]
Update KERNEL.SKYLAKEX
Martin Kroeker [Mon, 16 Mar 2020 20:59:55 +0000 (21:59 +0100)]
Merge pull request #2515 from zelong-1024/develop
[OpenBLAS]: benchmark for her/her2 LEVEL2 functions
Martin Kroeker [Mon, 16 Mar 2020 20:58:55 +0000 (21:58 +0100)]
Merge pull request #2513 from aaawuanjun/develop
[OpenBlas]: Add benchmark tpsv file and modify benchmark/Makefile
Martin Kroeker [Mon, 16 Mar 2020 20:58:34 +0000 (21:58 +0100)]
Merge pull request #2516 from wjc404/develop
AVX2 STRSM kernels
wjc404 [Mon, 16 Mar 2020 16:39:37 +0000 (16:39 +0000)]
Update KERNEL.ZEN
wjc404 [Mon, 16 Mar 2020 16:34:08 +0000 (00:34 +0800)]
AVX2 STRSM kernel
l00536773 [Mon, 16 Mar 2020 03:19:05 +0000 (11:19 +0800)]
[OpenBLAS]: benchmark for her/her2 LEVEL2 functions
[description]: benchmark for her/her2
[solution]: added benchmark for her/her2, modified makefile in benchmark
[dts]:
Martin Kroeker [Sat, 14 Mar 2020 13:21:30 +0000 (14:21 +0100)]
Merge pull request #2508 from liujingjue/develop
[OpenBLAS]:fix the iamax benchmark error
Martin Kroeker [Sat, 14 Mar 2020 12:27:40 +0000 (13:27 +0100)]
Merge pull request #2512 from martin-frbg/lapackh
Move declarations of lapack_complex_custom types outside the extern C
Martin Kroeker [Sat, 14 Mar 2020 12:08:36 +0000 (13:08 +0100)]
Merge pull request #2506 from xiaofengF/develop
Add benchmark for SPMV and fix segmentation fault when data size >= 50000
wuanjun 00447568 [Sat, 14 Mar 2020 01:11:08 +0000 (09:11 +0800)]
[OpenBlas]: Add benchmark tpsv file and modify benchmark/Makefile
[Description]: Solve lack of tpsv benchmark.
Martin Kroeker [Fri, 13 Mar 2020 22:16:35 +0000 (23:16 +0100)]
Merge pull request #2505 from aaawuanjun/develop
[OpenBlas]:Add benchmark tpmv.c and modify benchmark/Makefile
Martin Kroeker [Fri, 13 Mar 2020 22:04:01 +0000 (23:04 +0100)]
Merge pull request #2511 from martin-frbg/fixppctest
Prevent attempts to run ctest or test when fortran is not available
Martin Kroeker [Fri, 13 Mar 2020 19:34:13 +0000 (20:34 +0100)]
Move declarations of lapack_complex_custom types outside the extern C
fixes #2510
Martin Kroeker [Fri, 13 Mar 2020 19:11:19 +0000 (20:11 +0100)]
Do not attempt to run test without fortran
Martin Kroeker [Fri, 13 Mar 2020 19:10:26 +0000 (20:10 +0100)]
Do not attempt to run ctest without fortran
The main Makefile takes care of this in the build process, but users or CI jobs may try to run this directly
l00546269 [Fri, 13 Mar 2020 02:58:39 +0000 (10:58 +0800)]
[OpenBLAS]:fix the iamax benchmark error
[Description]:the result for i?amax is not MFlops, it is MBytes
jayfely@qq.com [Wed, 11 Mar 2020 09:02:34 +0000 (17:02 +0800)]
Remove cspmv and zspmv to remove the error occured in travis CI
jayfely@qq.com [Wed, 11 Mar 2020 08:36:45 +0000 (16:36 +0800)]
Modify Makefile in interface to remove the error occured in travis CI
jayfely@qq.com [Wed, 11 Mar 2020 07:48:58 +0000 (15:48 +0800)]
Only keep spmv.goto and spmv.atlas
wuanjun 00447568 [Wed, 11 Mar 2020 04:31:48 +0000 (12:31 +0800)]
[OpenBlas]:Add benchmark tpmv.c and modify Makefile
[Description]:Solve the problem of missing tpmv.c benchmark file
jayfely@qq.com [Wed, 11 Mar 2020 02:30:09 +0000 (10:30 +0800)]
Update spmv.c: solve segmentation fault when m and n are larger than 50000
Martin Kroeker [Tue, 10 Mar 2020 22:38:07 +0000 (23:38 +0100)]
Merge pull request #2503 from martin-frbg/xerbl
Apply fix for LAPACK issue 394 (fixed-form code beyond column 72)
Martin Kroeker [Tue, 10 Mar 2020 19:01:23 +0000 (20:01 +0100)]
Merge pull request #2502 from martin-frbg/issue2497
Fix INTERFACE64 not propagating to the fortran codes on ARMV8
Martin Kroeker [Tue, 10 Mar 2020 15:44:40 +0000 (16:44 +0100)]
Merge pull request #2501 from jijiwawa/Fix_mistakes
Fix pr #2487 error
s00527847 [Tue, 10 Mar 2020 23:26:06 +0000 (19:26 -0400)]
Use the correct unit of measure
Martin Kroeker [Tue, 10 Mar 2020 12:37:41 +0000 (13:37 +0100)]
Apply fix for Reference-LAPACK issue 394
reference to XERBLA extending beyond column 72, breaking builds with compilers that default to traditional punch card format
Martin Kroeker [Tue, 10 Mar 2020 11:51:07 +0000 (12:51 +0100)]
Restore INTERFACE64 for arm64
Martin Kroeker [Tue, 10 Mar 2020 11:49:21 +0000 (12:49 +0100)]
Merge pull request #37 from xianyi/develop
rebase
jayfely@qq.com [Tue, 10 Mar 2020 06:32:18 +0000 (14:32 +0800)]
Modify Makefile in Benchmark
jayfely@qq.com [Tue, 10 Mar 2020 06:22:18 +0000 (14:22 +0800)]
Add benchmark for SPMV
Zhang Xianyi [Mon, 9 Mar 2020 08:04:33 +0000 (16:04 +0800)]
Merge pull request #2498 from njutcz/develop
Add benchmark for ?amax, ?max, ?amin, ?min, i?max, i?amin and i?min.
s00548429 [Mon, 9 Mar 2020 07:36:50 +0000 (15:36 +0800)]
Fix the functional bugs for zamax.
s00548429 [Mon, 9 Mar 2020 06:59:03 +0000 (14:59 +0800)]
Add benchmark for ?amax, ?max, ?amin, ?min, i?max, i?amin and i?min.
njutcz [Mon, 9 Mar 2020 02:39:40 +0000 (10:39 +0800)]
Merge pull request #1 from xianyi/develop
update
Martin Kroeker [Sun, 8 Mar 2020 07:09:58 +0000 (08:09 +0100)]
Merge pull request #2495 from ZuoQ3/develop
add benchmark for axpby test
Martin Kroeker [Sat, 7 Mar 2020 22:04:21 +0000 (23:04 +0100)]
Merge pull request #2494 from shengyang-3390/develop
add benchmark for csrot and zdrot
Martin Kroeker [Sat, 7 Mar 2020 21:26:00 +0000 (22:26 +0100)]
Merge pull request #2489 from jijiwawa/brightness
Remove redundant code
s00527847 [Sat, 7 Mar 2020 18:09:19 +0000 (13:09 -0500)]
add trmm.c
s00527847 [Wed, 4 Mar 2020 22:44:50 +0000 (17:44 -0500)]
Remove redundant code